diff --git a/00-META/README.md b/00-META/README.md index 97d91c7..6e159c1 100644 --- a/00-META/README.md +++ b/00-META/README.md @@ -34,7 +34,7 @@ comparing it against the code rather than by anyone noticing: - It described the pipeline as having a separate builder process and a build stage that packages. Neither was true after 2026-08-04; the documents stayed stale until 2026-08-06 - ([ADR 0014](../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)). + ([ADR 0010](../02-DECISIONS/0010-delivery.md)). - It listed the mesh as spanning a fixed number of named machines, which is exactly the content this repository cannot carry. @@ -44,10 +44,10 @@ symlinks at all — the rule is not merely "only the installer may link", and a elevating linking to a principle points the opposite way from where this is going. What exists today is that the installer owns and reconciles every link -([ADR 0011](../02-DECISIONS/0011-the-installer-owns-linking.md)) — an as-is fact, recorded in +([ADR 0012](../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md)) — an as-is fact, recorded in [`03-DESIGN/00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md). Centralising who may link narrowed the incident class; it did not close it. The intent is to -remove the mechanism, recorded as [ADR 0018](../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md). +remove the mechanism, recorded as [ADR 0012](../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md). A founding document contradicting the direction of travel is precisely the failure this folder exists to prevent. diff --git a/00-META/checks/README.md b/00-META/checks/README.md new file mode 100644 index 0000000..7189f02 --- /dev/null +++ b/00-META/checks/README.md @@ -0,0 +1,50 @@ +# Checks + +``` +python3 00-META/checks/records.py structure: links, citations, supersession, topics +python3 00-META/checks/index.py the reading order in 02-DECISIONS/README.md is current +python3 00-META/checks/index.py --write regenerate it +``` + +Non-zero exit on any problem, so it can be a gate rather than a report. + +**Why this exists.** Until now nothing in this repository was verified by anything but reading, +which is how a superseded decision stayed live in the constitution for days and in +`01-to-be/README.md` alongside it. Both were found by a person looking. `how-we-build` §5 says +*an unenforced rule is indistinguishable from a wrong one, and costs more, because people +believe it* — this repository was carrying several. + +**Every check here failed on something real before it passed.** A check that has never failed is +indistinguishable from one that cannot. + +| Check | Asserts | Found | +|---|---|---| +| `links` | every relative link resolves | — (run ad hoc during authoring; now permanent) | +| `rests-on` | `decisions:` and `extends:` name records that exist and are **accepted** | the class behind both incidents | +| `live-citation` | a governing document citing a **superseded** record names its replacement in the same paragraph | `01-to-be/README.md` citing ADR 0022 as live guidance | +| `supersession` | if A says it was superseded by B, B says it supersedes A | ADR 0012 never declared that it superseded 0011 | +| `numbering` | the number in the filename is the number in the heading | — | +| `topics` | every record names a topic the index knows | — | +| *(index.py)* | the written reading order matches what the records say | — | +| `status-vs-code` | a to-be document naming specific code is not still `designed` | **ten documents**, several with a *What was built* section, describing lab-proven code | + +## What is deliberately not checked + +- **`02-DECISIONS/` and `01-RESEARCH/` may cite superseded records freely.** A decision record + discusses history; research records what was observed. Flagging those would produce noise on + correct documents, and a check that cries wolf gets suppressed — which costs more than not + having it. +- **`03-DESIGN/00-as-is/` may rest on a superseded record.** It describes what runs, and what + runs was built under whatever was decided at the time + ([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md): + *as-is describing a superseded decision is exactly what as-is is for*). +- **Whether a citation's prose is still true.** Only whether the record it points at is live. + A document can cite an accepted record and describe it wrongly, and nothing here notices. + +So "governing" means `00-META/` and `03-DESIGN/01-to-be/` — the documents that tell somebody +what to do. + +## Adding a check + +State what incident it would have caught, and make it fail before you make it pass. A check +whose failure has never been observed is a guess about its own correctness. diff --git a/00-META/checks/index.py b/00-META/checks/index.py new file mode 100644 index 0000000..22a2f80 --- /dev/null +++ b/00-META/checks/index.py @@ -0,0 +1,108 @@ +#!/usr/bin/env python3 +"""Generate the decision index, and check the written one still matches. + +A number identifies a record and never changes, so the folder listing is creation order rather +than reading order. The index is what carries the path — and it is written rather than only +generated on demand, because a reader on a forge sees the folder and not a command. + +The objection to a written index is that it drifts. That objection is answered by checking it +rather than by refusing to write one, which is `how-we-build` §5: a rule states how it is +checked. + + python3 00-META/checks/index.py --write regenerate it + python3 00-META/checks/index.py fail if it is stale +""" + +import glob +import io +import os +import re +import sys + +README = "02-DECISIONS/README.md" +START = "" +END = "" + +# The reading order. Topics a record may belong to, in the order somebody would learn the system. +TOPICS = [ + ("the mesh", "What the mesh is"), + ("the tiers", "Its tiers, from the bottom up"), + ("what runs on it", "What runs on them, and how it gets there"), + ("building it", "How it is built"), + ("checking it", "How it is checked"), + ("how we work", "How we work"), +] + + +def field(text, name): + m = re.search(r'^%s:\s*(.+)$' % name, text, re.M) + return m.group(1).strip() if m else None + + +def records(): + out = [] + for path in sorted(glob.glob('02-DECISIONS/0*.md')): + text = io.open(path, encoding='utf-8').read() + heading = re.search(r'^# \d+\.\s*(.+)$', text, re.M) + out.append({ + "file": os.path.basename(path), + "number": os.path.basename(path)[:4], + "title": heading.group(1).strip() if heading else "(no heading)", + "topic": field(text, "topic"), + "status": field(text, "status"), + }) + return out + + +def render(rs): + known = {t for t, _ in TOPICS} + lines = [START, ""] + for topic, label in TOPICS: + rows = [r for r in rs if r["topic"] == topic] + if not rows: + continue + lines.append("### %s" % label) + lines.append("") + for r in rows: + mark = "" if r["status"] == "accepted" else " *(%s)*" % r["status"] + lines.append("- **%s** — [%s](%s)%s" % (r["number"], r["title"], r["file"], mark)) + lines.append("") + stray = [r for r in rs if r["topic"] not in known] + if stray: + lines.append("### Unfiled") + lines.append("") + for r in stray: + lines.append("- **%s** — [%s](%s) — `topic:` is %r, which is not one of %s" % ( + r["number"], r["title"], r["file"], r["topic"], ", ".join(sorted(known)))) + lines.append("") + lines.append(END) + return "\n".join(lines) + + +def main(): + text = io.open(README, encoding='utf-8').read() + wanted = render(records()) + + if START not in text or END not in text: + print("index: %s has no index markers (%s / %s)" % (README, START, END)) + return 1 + + current = text[text.index(START):text.index(END) + len(END)] + if "--write" in sys.argv: + if current == wanted: + print("index: already current") + return 0 + io.open(README, 'w', encoding='utf-8').write(text.replace(current, wanted, 1)) + print("index: written") + return 0 + + if current != wanted: + print("index: %s is stale. Regenerate it:\n" + " python3 00-META/checks/index.py --write" % README) + return 1 + print("index: current") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/00-META/checks/records.py b/00-META/checks/records.py new file mode 100644 index 0000000..4469db8 --- /dev/null +++ b/00-META/checks/records.py @@ -0,0 +1,333 @@ +#!/usr/bin/env python3 +"""Structural checks over HQ's own records. + +Every check here exists because the thing it checks for actually happened. See README.md +for which incident is behind which check. Run from the repository root: + + python3 00-META/checks/records.py + +Exits non-zero if anything fails, so it can be a gate rather than a report. +""" + +import os +import re +import sys + +ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) + +# Where a citation is *guidance* rather than history. A document here tells somebody what to +# do, so a link to a superseded record is an instruction to follow a withdrawn decision. +# 02-DECISIONS and 01-RESEARCH are deliberately absent: they record what was decided and what +# was observed, and both legitimately discuss superseded records at length. +GOVERNING = ("00-META/", "03-DESIGN/01-to-be/") + +LINK = re.compile(r"\[[^\]]*\]\((?!https?:|mailto:)([^)]+)\)") +ADR_FILE = re.compile(r"^(\d{4})-") + + +def markdown_files(): + for base, dirs, files in os.walk(ROOT): + dirs[:] = [d for d in dirs if d not in (".git", ".claude")] + for name in sorted(files): + if name.endswith(".md"): + yield os.path.join(base, name) + + +def rel(path): + return os.path.relpath(path, ROOT) + + +def read(path): + with open(path, encoding="utf-8") as handle: + return handle.read() + + +def frontmatter(text): + """Minimal frontmatter reader — enough for the fields these checks use. + + Not a YAML parser on purpose: a dependency in a repository that has none, to read four + scalar fields and one list, would cost more than it returns. + """ + if not text.startswith("---\n"): + return {} + end = text.find("\n---", 4) + if end == -1: + return {} + fields, key = {}, None + for line in text[4:end].split("\n"): + item = re.match(r"^\s+-\s+(.*)$", line) + if item and key: + fields.setdefault(key, []).append(item.group(1).strip()) + continue + pair = re.match(r"^([A-Za-z_-]+):\s*(.*)$", line) + if pair: + key = pair.group(1) + value = pair.group(2).strip() + if value.startswith("[") and value.endswith("]"): + # inline list: `decisions: [a, b]`, `code: []` + inner = value[1:-1].strip() + fields[key] = [v.strip() for v in inner.split(",") if v.strip()] + elif value: + fields[key] = value + else: + fields[key] = [] + return fields + + +def load_records(): + """Every decision record, by its four-digit number.""" + records = {} + folder = os.path.join(ROOT, "02-DECISIONS") + for name in sorted(os.listdir(folder)): + match = ADR_FILE.match(name) + if not name.endswith(".md") or not match: + continue + path = os.path.join(folder, name) + text = read(path) + records[match.group(1)] = { + "number": match.group(1), + "name": name, + "path": path, + "text": text, + "front": frontmatter(text), + } + return records + + +class Failures: + def __init__(self): + self.items = [] + + def add(self, check, location, message): + self.items.append((check, location, message)) + + def report(self): + if not self.items: + print("records: all checks passed") + return 0 + by_check = {} + for check, location, message in self.items: + by_check.setdefault(check, []).append((location, message)) + for check in sorted(by_check): + print(f"\n{check} — {len(by_check[check])} problem(s)") + for location, message in by_check[check]: + print(f" {location}\n {message}") + print(f"\nrecords: {len(self.items)} problem(s)") + return 1 + + +def check_links(failures): + """Every relative link resolves to something that exists.""" + for path in markdown_files(): + folder = os.path.dirname(path) + for number, line in enumerate(read(path).split("\n"), 1): + for target in LINK.findall(line): + target = target.split("#")[0].strip() + if not target: + continue + if not os.path.exists(os.path.normpath(os.path.join(folder, target))): + failures.add("links", f"{rel(path)}:{number}", f"link does not resolve: {target}") + + +def check_rests_on(failures, records): + """A document's `decisions:` and a record's `extends:` must name a live record. + + These are the load-bearing citations: the document declares that it rests on that + decision. Resting on a withdrawn one is the defect this whole check set exists for. + """ + for path in markdown_files(): + front = frontmatter(read(path)) + cited = list(front.get("decisions", []) or []) + extends = front.get("extends") + if isinstance(extends, str) and extends: + cited.append(extends) + + for entry in cited: + match = ADR_FILE.match(os.path.basename(entry)) + if not match: + failures.add("rests-on", rel(path), f"not a decision record: {entry}") + continue + number = match.group(1) + if number not in records: + failures.add("rests-on", rel(path), f"no such record: {entry}") + continue + status = records[number]["front"].get("status") + if status != "accepted": + # A proposed record may extend another proposed one. Decisions are drafted in + # chains -- 0059 extends 0057 while both await review -- and refusing that would + # mean either drafting out of order or marking records accepted to satisfy a + # check, which is the failure this repository already made once. + if frontmatter(read(path)).get("status") == "proposed": + continue + # An as-is document describes what runs, and what runs was built under + # whatever was decided at the time. ADR 0056: "as-is describing a superseded + # decision is exactly what as-is is for." + if rel(path).startswith("03-DESIGN/00-as-is/"): + continue + # An extension that supersedes legitimately names what it replaced. + this = ADR_FILE.match(os.path.basename(path)) + supersedes = records[number]["front"].get("superseded-by", "") + if this and supersedes and os.path.basename(path) in str(supersedes): + continue + failures.add( + "rests-on", + rel(path), + f"rests on ADR {number}, which is '{status}' — a document may not rest on a " + f"record that is not accepted", + ) + + +def check_live_citations(failures, records): + """In a governing document, a link to a superseded record must name its replacement. + + The reader of a rule needs to know the rule was withdrawn, and needs somewhere to go. + Naming the superseder in the same paragraph is both, and it is what a person would + write anyway. + """ + superseded = { + number: record["front"].get("superseded-by", "") + for number, record in records.items() + if record["front"].get("status") == "superseded" + } + + for path in markdown_files(): + if not any(rel(path).startswith(prefix) for prefix in GOVERNING): + continue + text = read(path) + offset = 0 + for paragraph in text.split("\n\n"): + line_no = text[:offset].count("\n") + 1 + offset += len(paragraph) + 2 + targets = LINK.findall(paragraph) + for target in targets: + match = ADR_FILE.match(os.path.basename(target.split("#")[0])) + if not match or match.group(1) not in superseded: + continue + number = match.group(1) + replacement = os.path.basename(str(superseded[number])) + if not replacement: + failures.add( + "live-citation", + f"{rel(path)}:{line_no}", + f"cites superseded ADR {number}, which names no superseder", + ) + continue + if not any(replacement in t for t in targets): + failures.add( + "live-citation", + f"{rel(path)}:{line_no}", + f"cites superseded ADR {number} without naming its replacement " + f"({replacement}) in the same paragraph", + ) + + +def check_supersession_symmetry(failures, records): + """If A says it was superseded by B, B must say it supersedes A.""" + for number, record in records.items(): + front = record["front"] + status = front.get("status") + by = os.path.basename(str(front.get("superseded-by", ""))) + + if status == "superseded" and not by: + failures.add("supersession", rel(record["path"]), "marked superseded but names no superseder") + continue + if by and status != "superseded": + failures.add("supersession", rel(record["path"]), f"names a superseder but status is '{status}'") + if not by: + continue + + match = ADR_FILE.match(by) + if not match or match.group(1) not in records: + failures.add("supersession", rel(record["path"]), f"superseder does not exist: {by}") + continue + other = records[match.group(1)] + claims = os.path.basename(str(other["front"].get("supersedes", ""))) + if claims != record["name"]: + failures.add( + "supersession", + rel(other["path"]), + f"ADR {number} says this supersedes it; this record does not say so " + f"(supersedes: {claims or 'absent'})", + ) + + +def check_topics(failures, records): + """Every record names a topic the index knows. + + The topic is what puts a record in the reading order, so a record without one — or with one + nobody defined — disappears from the index rather than appearing in the wrong place. That is + the quiet failure, so it is the one checked. + """ + known = {"the mesh", "the tiers", "what runs on it", "building it", "checking it", + "how we work"} + for number, record in sorted(records.items()): + topic = record["front"].get("topic") + if not topic: + failures.add("topics", rel(record["path"]), + "no topic, so it has no place in the reading order") + elif topic not in known: + failures.add("topics", rel(record["path"]), + "topic %r is not one of: %s" % (topic, ", ".join(sorted(known)))) + + +def check_numbering(failures, records): + """The number in the filename is the number in the heading.""" + for number, record in records.items(): + heading = re.search(r"^# (\d+)\.", record["text"], re.M) + if not heading: + failures.add("numbering", rel(record["path"]), "no '# N. Title' heading") + elif heading.group(1) != number.lstrip("0"): + failures.add( + "numbering", + rel(record["path"]), + f"filename says {number}, heading says {heading.group(1)}", + ) + + +def check_status_against_code(failures): + """A design document naming specific code may not still call itself `designed`. + + **Naming a file is a claim that the file implements this**, so the two fields have to agree. + They drifted: ten to-be documents named working, lab-proven code — several with a *What was + built* or *Raised, and observed* section — while still saying nothing had been built. + + Deliberately weak, and that is the point of it being mechanical. It cannot tell whether the + prose is true, only that a document has stopped claiming to be unbuilt once it points at + something. `code: [mesh-control]` — a repository with no path — is a plan and stays + `designed`. + """ + for path in markdown_files(): + if not rel(path).startswith("03-DESIGN/01-to-be/") or path.endswith("README.md"): + continue + front = frontmatter(read(path)) + if front.get("status") != "designed": + continue + for entry in front.get("code") or []: + named = re.sub(r"\s*\(.*\)$", "", entry).strip().split(None, 1) + if len(named) > 1: + failures.add( + "status-vs-code", + rel(path), + f"`designed`, but names {named[1]!r} in {named[0]}. Naming a file claims " + f"it implements this — use `in-progress`, or `implemented` once it is " + f"defensible from that repository's main branch.", + ) + break + + +def main(): + failures = Failures() + records = load_records() + check_links(failures) + check_rests_on(failures, records) + check_live_citations(failures, records) + check_supersession_symmetry(failures, records) + check_numbering(failures, records) + check_topics(failures, records) + check_status_against_code(failures) + print(f"records: {len(records)} decision records checked") + return failures.report() + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/00-META/how-we-build.md b/00-META/how-we-build.md index 4f1c98f..53bd6d8 100644 --- a/00-META/how-we-build.md +++ b/00-META/how-we-build.md @@ -3,7 +3,7 @@ status: canonical updated: 2026-08-23 derives: knowledge-base constitution page decisions: - - 02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md + - 02-DECISIONS/0020-the-mesh-is-governed-by-a-constitution.md --- # How we build @@ -37,13 +37,13 @@ incident behind it is not written down, and the fix is to write it down, not to | Rule | What it means | |---|---| | **Never write to a production database directly** | No insert, update, delete or schema statement executed against production by hand. Schema changes go through numbered migrations; data changes go through application code or the module's own capabilities. Raw statements skip every side effect the proper path has — events, audit, cache invalidation, fan-out. | -| **Every schema change is a migration** | Numbered, in the module's own language, compiled with it. Both a baseline for a fresh installation *and* an incremental migration for installations that already exist. If code references a column, the migration creating it must exist. [ADR 0006](../02-DECISIONS/0006-schema-changes-are-numbered-migrations.md) | +| **Every schema change is a migration** | Numbered, in the module's own language, compiled with it. Both a baseline for a fresh installation *and* an incremental migration for installations that already exist. If code references a column, the migration creating it must exist. [ADR 0013](../02-DECISIONS/0013-schema-changes-are-numbered-migrations.md) | | **Never bypass the pipeline** | No manual database edit, no manual restart as a workaround. Fix the cause and deploy. A workaround that works is a workaround that is never removed, and the next person cannot tell the node from its declaration. | -| **Never create a symlink** | A hand-made link caused production data loss through container volume resolution, and the judgement needed to make a safe exception is exactly the judgement unavailable at the moment it matters. **The mesh creates none at all** ([ADR 0018](../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md), which supersedes [ADR 0011](../02-DECISIONS/0011-the-installer-owns-linking.md)). The links the installer still reconciles are a migration, not a permission. | +| **Never create a symlink** | A hand-made link caused production data loss through container volume resolution, and the judgement needed to make a safe exception is exactly the judgement unavailable at the moment it matters. **The mesh creates none at all** ([ADR 0012](../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md), which supersedes [ADR 0012](../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md)). The links the installer still reconciles are a migration, not a permission. | | **Never push directly to the main branch** | Branch, push, review, merge. Every merge is a human checkpoint, without exception — **including in this repository**. A documentation repository is not a lower tier of care; a decision record lands the same way a service does. | -| **One change per pull request, and never merge unapproved work** | Unrelated improvements bundled together cannot be reviewed or reverted separately. And the checkpoint is **a person deciding, not a person clicking** — work may be merged by whoever wrote it once a human has explicitly approved *that merge*, and never on a standing permission, an instruction to do the work, silence, or the author's own judgement that it is ready. [ADR 0042](../02-DECISIONS/0042-approval-is-the-checkpoint.md) | +| **One change per pull request, and never merge unapproved work** | Unrelated improvements bundled together cannot be reviewed or reverted separately. And the checkpoint is **a person deciding, not a person clicking** — work may be merged by whoever wrote it once a human has explicitly approved *that merge*, and never on a standing permission, an instruction to do the work, silence, or the author's own judgement that it is ready. [ADR 0023](../02-DECISIONS/0023-approval-is-the-checkpoint.md) | | **Never open a pull request unprompted** | A permissions list saying it is allowed is not a request. | -| **A failed step fails the job** | A sequence that continues past a failure does the next thing in the wrong place. Gate each step on the last. [ADR 0008](../02-DECISIONS/0008-a-failed-step-fails-the-job.md), and §5. | +| **A failed step fails the job** | A sequence that continues past a failure does the next thing in the wrong place. Gate each step on the last. [ADR 0010](../02-DECISIONS/0010-delivery.md), and §5. | ### A failed step must stop the steps after it — how it was earned @@ -69,7 +69,7 @@ reported failure, nothing stopped, and the damage happened somewhere nobody was - **Every runtime variable is declared.** A variable the module reads and the manifest does not declare is invisible to the mesh: it will not be generated, injected, or audited. - **Provisioned credentials arrive through declared requirements**, never hardcoded in code, - compose files or scripts. [ADR 0005](../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md) + compose files or scripts. [ADR 0009](../02-DECISIONS/0009-modules-and-the-graph.md) - **Never install a package by hand.** A package is declared in the manifest and arrives the way every other package does. A hand-installed package is invisible to the mesh: it is not declared, not reproduced on the next node, and not present after a rebuild — and the node @@ -82,7 +82,7 @@ reported failure, nothing stopped, and the damage happened somewhere nobody was - **Every standalone application gets its own repository**, with a manifest at its root, registered as a build source. Creating an application directory in the monorepo is a convention violation and reviewers reject it. - [ADR 0010](../02-DECISIONS/0010-applications-live-in-their-own-repository.md) + [ADR 0015](../02-DECISIONS/0015-applications-live-in-their-own-repository.md) ### Migrations @@ -96,7 +96,7 @@ reported failure, nothing stopped, and the damage happened somewhere nobody was surface is regenerated from the mesh database; a local edit survives one synchronisation and is then silently overwritten, bringing back whatever it fixed. Use the mesh operation that owns the value. If unsure whether a file is managed, ask the tooling — the answer is not visible -from the file. [ADR 0004](../02-DECISIONS/0004-managed-files-are-generated-never-edited.md) +from the file. [ADR 0011](../02-DECISIONS/0011-managed-files-are-generated-never-edited.md) --- @@ -114,12 +114,14 @@ and it runs on no node at all. Anatomy makes attractive names and poor boundaries. Name the thing the domain calls it. -### Group by domain, not by single function +### Things that change together share an authority, not a package -A module is a purpose, not a piece of software. Four modules that together constitute "how a -node is reachable" and cannot be assigned, versioned or replaced as one thing are four -accidents, not four boundaries. -[ADR 0017](../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) +When several modules always change together under one intent, name the **context** that decides +for them. Do not merge them into one module: they are delivered to different nodes, and a module +that must be assigned where half of it is unwanted is not a boundary either. + +Coherence is a context. Delivery is a module. Relationships are edges, not folders. +[ADR 0009](../02-DECISIONS/0009-modules-and-the-graph.md) ### Contexts integrate through the record, never through a shared schema @@ -224,7 +226,7 @@ No drive-by edits. Every change traces to a recorded decision. a design meeting with at least two node operators — which has never been met and cannot be, as there is one operator. A rule that cannot be satisfied is not a high standard; it is a rule everything silently violates. Recorded here as resolved in favour of what is achievable, and -what has in fact been practised ([ADR 0040](../02-DECISIONS/0040-the-constitution-absorbs-what-is-enforced.md)). +what has in fact been practised ([ADR 0022](../02-DECISIONS/0022-the-constitution-absorbs-what-is-enforced.md)). --- @@ -241,7 +243,7 @@ process. Absence of an override means these rules apply unmodified. ## 8. Code quality *Absorbed 2026-08-26 from the enforced page, which carried these rules while this document did -not — [ADR 0040](../02-DECISIONS/0040-the-constitution-absorbs-what-is-enforced.md).* +not — [ADR 0022](../02-DECISIONS/0022-the-constitution-absorbs-what-is-enforced.md).* **These rules are recorded because they are enforced, not because this repository earned them.** Every other rule here states the incident or measurement behind it. These state nothing, @@ -270,7 +272,7 @@ Data access, business logic and the interface layer are separate. *Scope: the mesh's services and surfaces. Tier 0 is a statically linked binary that must depend on nothing installed first, and is written in Go — -[ADR 0041](../02-DECISIONS/0041-the-host-depends-on-nothing.md).* +[ADR 0005](../02-DECISIONS/0005-the-node-host.md).* - TypeScript throughout; no new untyped JavaScript. - Strict, with no implicit `any` and no unchecked index access. diff --git a/00-META/process/00-overview.md b/00-META/process/00-overview.md index c89213d..ba54ecf 100644 --- a/00-META/process/00-overview.md +++ b/00-META/process/00-overview.md @@ -49,6 +49,8 @@ the expensive half. | [03](03-issues.md) | Issues | Something is wrong — often with the owner unknown | | [04](04-build-handoff.md) | Build handoff | A design is ready to be built | | [05](05-constitution-sync.md) | Constitution sync | `how-we-build.md` changed a rule the mesh enforces | +| [06](06-writing-a-module.md) | Writing a module | Something that runs today must run on the mesh | +| [07](07-feature-branches.md) | Feature branches across repos | Work that changes code, in one repo or several at once | ## Status lives in frontmatter diff --git a/00-META/process/06-writing-a-module.md b/00-META/process/06-writing-a-module.md new file mode 100644 index 0000000..9edadce --- /dev/null +++ b/00-META/process/06-writing-a-module.md @@ -0,0 +1,107 @@ +# Playbook 06 — Writing a module + +**Trigger.** Something that runs today must run on the mesh, or a new capability must be +declarable. + +**Who runs it.** Whoever is porting or writing it. + +*Written 2026-09-01 from doing this for the first time end to end. Every step below exists +because skipping it cost something.* + +## Before anything: read what runs + +**A module is written from the thing, not from memory of the thing.** For a port, that means its +current compose file, its environment, and where its data actually sits. Assumptions about any of +the three have been wrong every time they were not checked. + +Three questions, answered from the machine: + +| | why it decides something | +|---|---| +| **what containers, and how do they find each other?** | more than one means a `network`; names between them must match what the software is configured to dial | +| **where is its data?** | a bind mount moves with a path; a named volume does not; an anonymous volume is already losing data on every redeploy | +| **which values are secret, and which are merely settings?** | a secret goes in `own-secrets` or a grant; a setting goes in the manifest and may be overridden per node | + +## The steps + +1. **Name what it provides and requires**, if anything. A name is what a consumer is coupled to, + not the role it plays ([ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md)): + `postgres-database`, not `database`. Most modules provide nothing and require nothing — an + application is usually a leaf. + +2. **Declare capabilities, not dependencies, for facts about the machine.** `container-runtime`, + `package-manager`, `seat`. A capability is detected and refused against; it is not something a + module can install. + +3. **Write the resources in the order they must happen.** They are applied in the order written + and orphans are removed in reverse, so a `network` is written before the containers that join + it and removed after them. + +4. **Put every secret in a file, never in `env`.** A declaration travels over the broker in plain + text: a password in `env` is a password the broker sees. **The mesh delivers parts; a module + that needs them combined combines them.** + + **A sealed file holds the password and nothing else** — no key, no `=`, no newline that means + anything. So `env-file` must never point at one. It points at a file the module *declares*, + whose content leaves a hole: + + ``` + own-secrets superuser → /var/lib/postgres/superuser.secret the password, alone + a file /var/lib/postgres/superuser.env, mode 0600, + content: POSTGRES_PASSWORD=${secret:superuser} + the container env-file: [/var/lib/postgres/superuser.env] + ``` + + The host fills the hole on the machine, which is the only place both halves exist — the mesh + discarded the value + ([credentials and their rotation](../../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)). + A **provisioner** is the exception: it reads a password file, so it mounts the `.secret` + directly. + + Every example module in `mesh-control` had this wrong and shipped: `own-secrets` pointing at a + path *named* `.env`, mounted as `env-file`, holding a bare password. Docker reads that as a + malformed line and the container starts **with no password set at all** — not a failure to + start, a service running on the wrong credential. They parsed and they resolved. Two tests in + `examples/modules` now refuse both halves of it. + + Add `restart-on` naming the env file, or the container keeps the credential it started with + through every rotation. + +5. **Pin every image by digest.** A tag moves. The manifest in a repository names artifacts; the + manifest the mesh holds names digests, and they are not the same document. + +6. **Decide generate or accept.** A new module's credential is generated. **An adopted one keeps + the credential it already has** — `secret accept` — because minting a new password for a + database that already exists locks the application out of its own data. + +7. **Add a provisioner only if the software cannot read a file.** A proxy that watches a + directory needs nothing. PostgreSQL needs `CREATE ROLE`, an object store needs a bucket and a + policy, an identity provider needs a realm and a client — those need a small program beside + them. It reads what the mesh granted and reconciles; it does not decide anything. + +8. **Prove it in the lab, against the real software.** Not that a container started — that the + thing works: the credential authenticates, a wrong one is refused, the containers reach each + other, the data survives a restart. + +## What the first port actually cost + +Six attempts, one real bug. Recorded because the ratio is the lesson: **the mesh was right every +time and the scaffolding was not.** + +- A shape existed in the language and no host implemented it, so every declaration carrying one + was refused whole — correctly, and the host said exactly that. **Nobody was reading the host's + log.** Read it first; it is the only place that says why a machine did nothing. +- A blind find-and-replace renamed a provision in quotes and missed the same word bare. +- A command was tested only for the invocations that should fail, so it rejected every real one + and the suite stayed green. +- A test asserted on a helper rather than on the code that calls it, three separate times. **A + test that cannot fail when the behaviour is deleted is not defending the behaviour.** + +## Rules + +- **Read the host's log before theorising.** A declaration that was sent and not applied says so + there and nowhere else. +- **A failing test is kept, not skipped.** It is the reproduction. +- **Never rotate during an adoption.** Rotation is a separate act, afterwards, deliberately. +- **A data directory is never removed by the mesh** ([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)), + and that protects against the mesh only — not against a disk or a mistaken command. diff --git a/00-META/process/07-feature-branches.md b/00-META/process/07-feature-branches.md new file mode 100644 index 0000000..bb5f451 --- /dev/null +++ b/00-META/process/07-feature-branches.md @@ -0,0 +1,57 @@ +# Playbook 07 — Feature branches across repos + +**Trigger.** Work that changes code — in one code repo or in several at once (`mesh-sdk`, +`mesh-control`, `mesh-catalog`, `mesh-host`, `mesh-lab`, and `hq` when a decision rides along). + +**Who runs it.** Anyone who writes code, engineers and agents alike. Agents follow it exactly — +it is the guard against the failure it was written for. + +## The failure it prevents + +A feature was worked as a branch-and-MR per *unit of thought* — one per decision, one per +stacked increment — and each MR was treated as finished when it was *opened*, not when it was +*merged*. Across repos the same feature took a different branch name in each. The MRs piled up +unmerged: one session left **sixteen** stacked intermediate MRs that had to be consolidated and +closed by hand. An MR is a review checkpoint, not a scratchpad. + +## The rule + +One feature is **one branch name**, **one worktree per repo**, **one MR per repo**, opened +**once, at the end**. + +1. **Name the feature once.** `feat/`. The *same* branch name in every repo the feature + touches — never a different name per repo, never a fresh branch per increment within the + feature. +2. **Isolate each repo.** One git worktree per touched repo under `.work//`, branched + off `main`: + ``` + git worktree add .work// -b feat/ origin/main + ``` + Parallel features never collide, and no shared checkout is edited. +3. **Commit as you go — locally.** Increments land on the one branch. Nothing is pushed and no + MR is opened mid-feature. +4. **Finish, then publish.** When the whole feature is done — every repo, tests green — push + every branch and open **one MR per touched repo**, together. +5. **Merge promptly, once approved.** Every merge into `main` is notified and approved + ([ADR 0023](../../02-DECISIONS/0023-approval-is-the-checkpoint.md)); once it is, merge — + do not leave it sitting. The branch is deleted on merge. +6. **Leave nothing behind.** After the MRs merge, no `feat/` branch and no `.work/` + worktree survive. + +## What this is not + +- **Not a licence to batch unbounded work.** A feature is a *bounded* unit; if it sprawls for + days, end-of-feature bloat merely replaces per-increment bloat. Split it into features, each + its own branch and MR. +- **Not a second trunk.** Every repo branches off `main`. There is no longer an `initialization` + trunk. + +## How it is checked + +The end state is visible, and its absence is the smell: + +- After a feature merges, `git branch -r | grep feat/` and `git worktree list` return + nothing for it. A surviving branch or worktree means step 6 was skipped. +- More than one open MR in a repo that share no feature name, or a stack of MRs none of which is + merged, is the failure this playbook exists to prevent — stop and consolidate before opening + more. diff --git a/00-META/repos.md b/00-META/repos.md index 87e954b..d8e9194 100644 --- a/00-META/repos.md +++ b/00-META/repos.md @@ -15,26 +15,26 @@ and a forge address is an operational detail (see [`README`](../README.md)). | Repository | Owns | |---|---| | `hal` | The monorepo — the node runtime, the module catalogue, the delivery machinery, and the bootstrap scripts. Every core module lives here. | -| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0028](../02-DECISIONS/0028-hq-is-company-scoped.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. | +| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. | | *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. | ## What the mesh becomes -[ADR 0030](../02-DECISIONS/0030-the-repository-structure.md) records the repositories the -monorepo decomposes into. **`mesh-lab` and `mesh-host` exist so far** — the lab is built first -([ADR 0029](../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)); the rest are the +[ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md) records the repositories the +monorepo decomposes into. **`mesh-lab`, `mesh-host` and `mesh-control` exist so far** — the lab is built first +([ADR 0016](../02-DECISIONS/0016-the-lab.md)); the rest are the target, not the present. | Repository | Tier | Holds | |---|---|---| -| `mesh-host` | 0 | **exists.** The node host — one statically linked binary, requiring nothing present ([ADR 0041](../02-DECISIONS/0041-the-host-depends-on-nothing.md)) | +| `mesh-host` | 0 | **exists.** The node host — one statically linked binary, requiring nothing present ([ADR 0005](../02-DECISIONS/0005-the-node-host.md)) | | `mesh-substrate` | 1 | the four pinned services, as declarations | -| `mesh-control` | 2 | the control plane and its contexts | +| `mesh-control` | 2 | **exists.** The control plane and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) | | `mesh-surfaces` | 3 | tools, web, cli | -| `mesh-sdk` | — | contracts shared across tiers | +| `mesh-sdk` | — | the stable spine modules build against — the tool-serving harness, the messaging/event framework, the contracts and core primitives. Holds nothing per-module and nothing volatile ([ADR 0039](../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)). | | `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. | -Tier 4's shape is open, and deliberately so: see ADR 0030 and +Tier 4's shape is open, and deliberately so: see ADR 0019 and [research 005](../01-RESEARCH/005-domain-grouping/00-overview.md). ## What lives where inside the monorepo @@ -53,7 +53,7 @@ Named by role, because the layout is itself part of the as-is design — see ## Why applications do not live in the monorepo A standalone application in the monorepo is a convention violation, and reviewers reject it. -The reasoning is recorded in [`02-DECISIONS/0010`](../02-DECISIONS/0010-applications-live-in-their-own-repository.md): +The reasoning is recorded in [`02-DECISIONS/0010`](../02-DECISIONS/0015-applications-live-in-their-own-repository.md): the mesh installs, provisions for, and ships an application through exactly the same machinery whether or not its source sits beside the mesh's own — so co-location buys nothing and costs the monorepo's review cadence. @@ -64,7 +64,7 @@ Each module is a standalone package that consumes its dependencies from the priv not from a sibling directory. The workspace was removed after it caused build-versus-development divergence — a workspace member importing another resolved to local unbuilt source in the pipeline and to a published version in development. Recorded in -[`02-DECISIONS/0007`](../02-DECISIONS/0007-no-npm-workspace.md). +[`02-DECISIONS/0007`](../02-DECISIONS/0014-no-npm-workspace.md). Consequence, and it is a real one: a cross-package change is two steps — publish, then consume — and a repository-wide `npm install` does not exist. diff --git a/01-RESEARCH/001-module-domain-decomposition/00-overview.md b/01-RESEARCH/001-module-domain-decomposition/00-overview.md index ff6a9c0..20210a1 100644 --- a/01-RESEARCH/001-module-domain-decomposition/00-overview.md +++ b/01-RESEARCH/001-module-domain-decomposition/00-overview.md @@ -2,7 +2,7 @@ status: active initiated: 2026-08-22 touches: [03-DESIGN/00-as-is/02-modules-and-manifests.md, 03-DESIGN/00-as-is/10-module-catalogue.md, 03-DESIGN/01-to-be/00-work-breakdown.md] -became: [02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md] +became: [02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md] --- # 001 — Module domain decomposition @@ -54,7 +54,7 @@ Tracked in [`analysis.md`](analysis.md) under "Open questions". ## Deliberately not decided Recorded so they are not mistaken for oversights. Each is open, and each comes out of -[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md); this effort stays +[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md); this effort stays `active` until they are answered. | Question | Status | @@ -64,4 +64,4 @@ Recorded so they are not mistaken for oversights. Each is open, and each comes o | Catalogue destination — one repository or many. | Open. Phase 4. | | What the shared library keeps after extraction. | Open. Phase 3. | | Where human agent modality is recorded — which user, on which node, a human agent acts as. | Open. Required by the model; not yet stored. | -| Which domains the modules outside the platform core group into. | Open, from [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md), which settles the principle and deliberately not the list. | +| Which domains the modules outside the platform core group into. | Open, from [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md), which settles the principle and deliberately not the list. | diff --git a/01-RESEARCH/001-module-domain-decomposition/analysis.md b/01-RESEARCH/001-module-domain-decomposition/analysis.md index d7c1f5d..8050778 100644 --- a/01-RESEARCH/001-module-domain-decomposition/analysis.md +++ b/01-RESEARCH/001-module-domain-decomposition/analysis.md @@ -237,6 +237,6 @@ returns. That mechanism is the subject of a separate ADR. pipeline resolves dependencies across the registry rather than the filesystem? 4. **SDK residue** — after extraction, does `hal/sdk` keep transport (`amqp-client`), or does that belong to `hal/stream`? Everything imports it, which argues both ways. -5. **Human agent modality.** ADR 0015 requires a fact the mesh does not record: which +5. **Human agent modality.** ADR 0001 requires a fact the mesh does not record: which user, on which node, a human agent acts as. Where does it live — an attribute of the agent, or of the agent-node binding? diff --git a/01-RESEARCH/002-local-mesh/00-overview.md b/01-RESEARCH/002-local-mesh/00-overview.md index b0d5da5..26aad11 100644 --- a/01-RESEARCH/002-local-mesh/00-overview.md +++ b/01-RESEARCH/002-local-mesh/00-overview.md @@ -2,7 +2,7 @@ status: graduated initiated: 2026-08-22 touches: [03-DESIGN/00-as-is/04-delivery.md, 03-DESIGN/00-as-is/05-runtime-and-installation.md] -became: [03-DESIGN/01-to-be/01-end-to-end-testing.md, 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md] +became: [03-DESIGN/01-to-be/01-end-to-end-testing.md, 02-DECISIONS/0016-the-lab.md] --- # 002 — A mesh that runs locally @@ -24,7 +24,7 @@ This effort establishes what already runs in a container, what is welded to the what it would take to close the gap. It does **not** choose an approach: the central question — how a containerised node executes a module service, when a module service is defined today as a systemd unit shelling to `docker compose` in `/services/` — is not -answered by ADR 0015 and is recorded below rather than decided. +answered by ADR 0001 and is recorded below rather than decided. ## What was established diff --git a/01-RESEARCH/002-local-mesh/analysis.md b/01-RESEARCH/002-local-mesh/analysis.md index 785df6e..c14ca14 100644 --- a/01-RESEARCH/002-local-mesh/analysis.md +++ b/01-RESEARCH/002-local-mesh/analysis.md @@ -253,7 +253,7 @@ over either way. ## References -- [`02-DECISIONS/0001`](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) — the decision this +- [`02-DECISIONS/0001`](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) — the decision this phase unblocks - [`03-DESIGN/00-work-breakdown.md`](../../03-DESIGN/01-to-be/00-work-breakdown.md) — Phase 0 tasks and checkpoint diff --git a/01-RESEARCH/003-service-supervision/00-overview.md b/01-RESEARCH/003-service-supervision/00-overview.md index 4736fc6..a785cb1 100644 --- a/01-RESEARCH/003-service-supervision/00-overview.md +++ b/01-RESEARCH/003-service-supervision/00-overview.md @@ -1,8 +1,11 @@ --- -status: active +status: graduated initiated: 2026-08-22 touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md] -became: [] +became: + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0005-the-node-host.md + - 03-DESIGN/01-to-be/05-the-node-host.md --- # 003 — Who supervises a service @@ -40,14 +43,38 @@ This effort answers the cost half. It does not choose. - **There is a third option neither of us named**, and it is the one that also solves Phase 0: run HAL's own daemons as containers, making Docker the supervisor for everything. Local and production then have the same shape rather than a translation layer between them. -- **It cannot be all-or-nothing**, and ADR 0015 already says why: a human agent acts through a +- **It cannot be all-or-nothing**, and ADR 0001 already says why: a human agent acts through a shell and a desktop. Those parts are on the host by definition. - One incidental finding: the automatic node rescue that documentation describes **does not exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer. Detail and costs in [`analysis.md`](analysis.md). -## Decision needed +## What it became -Which supervision model the mesh adopts, recorded in a decision record before Phase 0 builds -anything. The options and their costs are in `analysis.md` under "Options". +*Closed 2026-08-28.* The decision this effort asked for was taken — and taken without citing it, +which is why the effort sat `active` for five days after being answered. Recorded here because +finding that is the point of a sweep. + +**The third option is what the mesh adopted.** `Docker is the supervisor for everything` is +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md): the +host is a plain process on the machine and everything above tier 0 is a container. The substrate +bootstrap declares no service at all — it is package, container, action, container — so the +44-of-44 restart policies this effort counted are the supervision, exactly as it argued. + +**Fate-sharing was the hard part, and it is solved the way this effort predicted.** It said any +mesh-native supervisor inherits the problem *unless it sits outside the mesh's own process +tree*. [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts +the launcher there: it supervises the host as a child and shares no code with it, so a host that +cannot start is still recovered. + +**It is not all-or-nothing, as this effort insisted.** A human agent acts through a shell and a +desktop, and those are on the host. So is the host itself — the one thing an init starts. + +## What is not closed + +**The automatic node rescue the documentation describes does not exist.** No unit declares +`OnFailure=`, and nothing calls the rescue script on a timer. That is a documented behaviour +which never happens, and it outlives this effort — filed as +[`04-ISSUES/008`](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md) +rather than closed with it. diff --git a/01-RESEARCH/003-service-supervision/analysis.md b/01-RESEARCH/003-service-supervision/analysis.md index b8f03d1..9a5aafb 100644 --- a/01-RESEARCH/003-service-supervision/analysis.md +++ b/01-RESEARCH/003-service-supervision/analysis.md @@ -128,10 +128,10 @@ What it costs, honestly: --- -## 5. Why it cannot be all-or-nothing — and ADR 0015 already says so +## 5. Why it cannot be all-or-nothing — and ADR 0001 already says so Some of what runs under systemd today **cannot** be containerised, and the reason is -already in the domain model. ADR 0015: +already in the domain model. ADR 0001: > a non-human agent acts through a spawned session — a human agent acts through a shell or > desktop @@ -223,7 +223,7 @@ fate-sharing reason in §3. ## References - [`002-local-mesh`](../002-local-mesh/analysis.md) — the effort this came out of -- [`02-DECISIONS/0001`](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) — agent modality, which +- [`02-DECISIONS/0001`](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) — agent modality, which decides what cannot leave the host - `modules/hal/meshware/daemon/src/cerebellum.ts:815-828` — the self-restart workaround - `modules/hal/meshware/systemd/hal-module@.service` — the per-module Docker lifecycle diff --git a/01-RESEARCH/004-lab-network/00-overview.md b/01-RESEARCH/004-lab-network/00-overview.md index 5a70246..1837334 100644 --- a/01-RESEARCH/004-lab-network/00-overview.md +++ b/01-RESEARCH/004-lab-network/00-overview.md @@ -3,9 +3,9 @@ status: graduated initiated: 2026-08-22 touches: [03-DESIGN/00-as-is/01-mesh-and-transport.md, 03-DESIGN/01-to-be/01-end-to-end-testing.md] became: - - 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md - - 02-DECISIONS/0031-the-lab-provides-the-underlay.md - - 02-DECISIONS/0033-a-router-is-scenery-not-a-node.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md - 03-DESIGN/01-to-be/02-scenario-declaration.md - 03-DESIGN/01-to-be/01-end-to-end-testing.md --- diff --git a/01-RESEARCH/005-domain-grouping/00-overview.md b/01-RESEARCH/005-domain-grouping/00-overview.md index 7a12a48..462badf 100644 --- a/01-RESEARCH/005-domain-grouping/00-overview.md +++ b/01-RESEARCH/005-domain-grouping/00-overview.md @@ -1,16 +1,18 @@ --- -status: active +status: graduated initiated: 2026-08-23 touches: - - 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md + - 02-DECISIONS/0009-modules-and-the-graph.md - 03-DESIGN/00-as-is/10-module-catalogue.md - 03-DESIGN/00-as-is/02-modules-and-manifests.md -became: [] +became: + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0009-modules-and-the-graph.md --- # 005 — Which domains the catalogue groups into -[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) settles +[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) settles that modules outside the platform core are grouped by domain rather than by single function, and deliberately does not settle the list. This effort settles the list — and, first, tests whether the premise survives measurement. @@ -23,7 +25,7 @@ together**, measured across the full history of the code repository. ## Why -The argument in ADR 0017 is that the catalogue's shape records what was installed rather than +The argument in ADR 0022 is that the catalogue's shape records what was installed rather than what anything is for — that four modules constituting "how a node is reachable" have no relationship the mesh can see, so a change to connectivity is made four times. @@ -51,7 +53,31 @@ premise**, in a way that narrows the effort usefully: The remaining work is the list itself, for the modules where grouping is justified, plus the open questions below. -## Open questions +## What it became + +*Closed 2026-08-28.* Three of the four questions are answered, and by records that did not cite +this effort — which is why it stayed open after being resolved. + +**Whether provider modules group at all** — *no.* +[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md): +there is no `networking` thing to install, there are concrete modules named individually. Folders +assert relationships; edges record them. *Provider* stops being a category at the same time. + +**Whether "group or leave" is even the right pair of options** — *it was not*, and that is the +useful finding. [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) +reframes it: things that change together share an **authority**, not a package. This effort's own +measurement is what that record rests on — reachability being the *only* place modules genuinely +co-change is why connectivity is a context and why nothing else needed one. + +**What to do with the ~50 modules that co-change with nothing** — *nothing.* They are modules. +Grouping is a tag and a query over the graph, neither of which anybody keeps true by hand. + +**What remains is a sequencing question, not a grouping one**, and it moves rather than closes: +*whether applications leave the monorepo before or after they group* is +[research 009](../009-migration/00-overview.md)'s, because it is about how to get from here to +there rather than about what the shape is. + +## Open (superseded by the above) questions | Question | Why it is open | |---|---| diff --git a/01-RESEARCH/005-domain-grouping/analysis.md b/01-RESEARCH/005-domain-grouping/analysis.md index 62b5f97..6070c51 100644 --- a/01-RESEARCH/005-domain-grouping/analysis.md +++ b/01-RESEARCH/005-domain-grouping/analysis.md @@ -10,7 +10,7 @@ updated: 2026-08-23 Every commit in the code repository's main branch that touches the module catalogue, reduced to the set of modules it touched. Platform-namespace modules are excluded — their decomposition is settled by -[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md). Modules that no +[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md). Modules that no longer exist are excluded, because pre-rename names dominate the raw signal and describe a catalogue nobody works in. @@ -95,14 +95,14 @@ remainder are genuine: | 2026-08-06 | firewall mesh-only by default, public by declaration | Each is one intent — *change how a node is reachable* — landing across the proxy, the -resolver, the firewall and the VPN together. That is exactly the shape ADR 0017 describes, and +resolver, the firewall and the VPN together. That is exactly the shape ADR 0022 describes, and it is the only place in the catalogue where the measurement finds it. The 2026-08-23 scoping commit is the sharpest case: it spans the reachability cluster **and** two providers, because "which network is this exposed on" is a reachability question asked of a database. -## What this means for ADR 0017 +## What this means for ADR 0022 The record's principle stands, and its scope needs narrowing. Grouping by domain is: @@ -130,7 +130,7 @@ Asked directly, and stated as an opinion because it is not yet decided. is implementation selection, a substantially larger design with its own failure modes, and nothing currently asks for it. 3. **It would hide which implementation serves a requirement** — the one place the mesh most - needs to be explicit, and precisely the indirection ADR 0017 warns grouping causes. + needs to be explicit, and precisely the indirection ADR 0022 warns grouping causes. A provider module is already exactly one purpose: it provisions one resource type. That is a boundary, not an accident of installation. diff --git a/01-RESEARCH/006-mesh-from-scratch/00-overview.md b/01-RESEARCH/006-mesh-from-scratch/00-overview.md index 653dfc9..176736b 100644 --- a/01-RESEARCH/006-mesh-from-scratch/00-overview.md +++ b/01-RESEARCH/006-mesh-from-scratch/00-overview.md @@ -2,8 +2,8 @@ status: active initiated: 2026-08-23 touches: - - 02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md - - 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md + - 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md + - 02-DECISIONS/0009-modules-and-the-graph.md - 03-DESIGN/00-as-is/00-overview.md - 03-DESIGN/01-to-be/00-work-breakdown.md became: [] @@ -24,9 +24,9 @@ disk, and where today's catalogue lands. ## Why Every structural decision so far has been a **correction**: eight contexts replacing thirty-three -modules ([ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)), domains +modules ([ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md)), domains replacing single-function modules -([ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md)). A +([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)). A correction inherits the frame of the thing it corrects, and two of the mesh's oldest problems look unsolvable from inside that frame: @@ -68,7 +68,7 @@ against taste: circle, self-hosted. Personal cloud infrastructure. 8. **Agents make it self-improving and self-healing.** 9. It is **end-to-end testable on one machine** - ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)). + ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). ## Status @@ -76,7 +76,7 @@ A first skeleton exists, with four design moves that the current shape does not `active` because two of them are unproven and one contradicts a record that is already accepted. -**Finding worth stating up front:** [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) +**Finding worth stating up front:** [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) names nine bounded contexts and **none of them owns connectivity** — no overlay, no resolution, no firewall, no ingress. Requirement 4 has no home in the accepted decomposition, while [research 005](../005-domain-grouping/analysis.md) found reachability to be the *only* part of @@ -87,8 +87,8 @@ the catalogue where modules genuinely change together under one intent. The skel | Question | Why it is open | |---|---| | Does the record — the event log contexts integrate through — belong to the substrate or the control plane? | It is infrastructure by shape and domain by content. Placing it wrong reintroduces a circularity. | -| One repository per tier, or per context? | Already open from ADR 0015 as "catalogue destination — one repository or many". The skeleton assumes per tier and does not settle it. | -| ~~Does an unprivileged node earn a place in the inventory, or only a presence?~~ | **Answered 2026-08-25** by the operator: a node is a *managed machine inside the mesh*, not an unprivileged something — and a disconnected node is still a node, in a different situation. The question posed a class distinction; the answer is that there is none, and what varies is **state**. Recorded as [ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md). | -| ~~Does absorbing overlay, filtering, packages, supervision and the container runtime make the host too large?~~ | **Answered 2026-08-25** — [`host-size.md`](host-size.md). Measured: the absorption is smaller than the machinery that already applies state, and eight of ten adapters already carry no dependency. The risk is not size but direction, and it is two modules wide. The claim survives with its scope corrected — the host carries one concern, *apply declared state on this machine*, of which the six are instances. Recorded as [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md), designed in [`05-the-node-host.md`](../../03-DESIGN/01-to-be/05-the-node-host.md). | -| Four substrate services or five? | The identity provider passes the tier test only if the control plane delegates authentication rather than doing it natively. | +| One repository per tier, or per context? | Already open from ADR 0001 as "catalogue destination — one repository or many". The skeleton assumes per tier and does not settle it. | +| ~~Does an unprivileged node earn a place in the inventory, or only a presence?~~ | **Answered 2026-08-25** by the operator: a node is a *managed machine inside the mesh*, not an unprivileged something — and a disconnected node is still a node, in a different situation. The question posed a class distinction; the answer is that there is none, and what varies is **state**. Recorded as [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md). | +| ~~Does absorbing overlay, filtering, packages, supervision and the container runtime make the host too large?~~ | **Answered 2026-08-25** — [`host-size.md`](host-size.md). Measured: the absorption is smaller than the machinery that already applies state, and eight of ten adapters already carry no dependency. The risk is not size but direction, and it is two modules wide. The claim survives with its scope corrected — the host carries one concern, *apply declared state on this machine*, of which the six are instances. Recorded as [ADR 0005](../../02-DECISIONS/0005-the-node-host.md), designed in [`05-the-node-host.md`](../../03-DESIGN/01-to-be/05-the-node-host.md). | +| ~~Four substrate services or five?~~ | **Answered conditionally**, which is the honest form — [`07-the-substrate.md`](../../03-DESIGN/01-to-be/07-the-substrate.md). The substrate is *what the control plane consumes and cannot grant itself*. The identity provider qualifies only if the control plane delegates authentication; if it authenticates natively it is an ordinary hosted service. The count follows from a decision not yet taken, and asserting four was asserting that decision. | | Does `feature` survive? | The skeleton splits it in two and argues the conflation is what makes the delivery pipeline hard to reason about. Unproven. | diff --git a/01-RESEARCH/006-mesh-from-scratch/code-skeleton.md b/01-RESEARCH/006-mesh-from-scratch/code-skeleton.md index 87398b4..62b1c42 100644 --- a/01-RESEARCH/006-mesh-from-scratch/code-skeleton.md +++ b/01-RESEARCH/006-mesh-from-scratch/code-skeleton.md @@ -160,8 +160,8 @@ neither option covers, and it is the most common one. | **Absorbed into the host** | It is not a module at all. It is part of what "managing a machine" means, and belongs in tier 0. | overlay membership, packet filtering, package management, service supervision, container runtime, filesystem management | | **Substrate** | The control plane cannot exist without it. Pinned, host-applied. | relational store, bus, object store, image registry | | **Control-plane context** | It decides something across nodes. | connectivity policy, inventory, delivery, provisioning, observability | -| **Workload module** | The mesh hosts it. Grouped per [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md). | media library, desktop session, collaboration tooling | -| **Leaves the repository** | A standalone application, per [ADR 0010](../../02-DECISIONS/0010-applications-live-in-their-own-repository.md). | the applications identified in research 005 | +| **Workload module** | The mesh hosts it. Grouped per [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md). | media library, desktop session, collaboration tooling | +| **Leaves the repository** | A standalone application, per [ADR 0015](../../02-DECISIONS/0015-applications-live-in-their-own-repository.md). | the applications identified in research 005 | **The first fate is the finding.** Research 005 measured the reachability cluster — proxy, resolver, firewall, overlay — as the only place in the catalogue where modules genuinely change @@ -256,7 +256,7 @@ mesh-surfaces/ TIER 3 mesh-catalog/ TIER 4 // layout as above -mesh-lab/ mesh-sdk/ hq/ (company-scoped — ADR 0028) +mesh-lab/ mesh-sdk/ hq/ (company-scoped — ADR 0019) ``` ## What this does not settle diff --git a/01-RESEARCH/006-mesh-from-scratch/skeleton.md b/01-RESEARCH/006-mesh-from-scratch/skeleton.md index ecd6868..502d5ec 100644 --- a/01-RESEARCH/006-mesh-from-scratch/skeleton.md +++ b/01-RESEARCH/006-mesh-from-scratch/skeleton.md @@ -80,12 +80,12 @@ mesh-surfaces/ TIER 3 — thin; no logic lives here cli/ the shell-facing interface mesh-catalog/ TIER 4 — what the mesh hosts - / grouped per ADR 0017, list per research 005 + / grouped per ADR 0022, list per research 005 mesh-lab/ the whole mesh, disposable, on one machine mesh-sdk/ contracts shared across tiers — types, not behaviour -hq/ company-scoped, not a mesh repository — ADR 0028 +hq/ company-scoped, not a mesh repository — ADR 0019 ``` ## The dependency rule @@ -97,7 +97,7 @@ a second surface would have to reimplement. This is the whole of the bootstrap answer, and per this repository's own rule it must say how it is checked: a dependency-direction lint in the build, failing on an upward import. A tier rule enforced by intention is the same as no tier rule — that is -[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied to architecture. +[ADR 0010](../../02-DECISIONS/0010-delivery.md) applied to architecture. ## Move 1 — the substrate is applied, not delivered @@ -144,7 +144,7 @@ assumption that every node is equivalent — already false, and today handled by ## Move 3 — connectivity becomes a context -[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) names nine contexts +[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) names nine contexts and none of them owns the overlay, the resolver, the firewall or the ingress. `config` owns PKI, which is the closest thing, and it is not close. @@ -162,8 +162,8 @@ So the evidence and the gap point the same way. `connectivity` owns: - certificates for both name spaces This is an addition to an accepted record, so it is a decision, not a drafting choice. It -belongs in a new record that extends ADR 0015 the way -[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) does — +belongs in a new record that extends ADR 0001 the way +[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) does — not written here. ### Where the networking actually lives @@ -195,7 +195,7 @@ The invitation was to check whether the concept survives. It does not, in one pi Today a **feature** means both *a thing built once* and *a thing selected per node*, and the delivery pipeline is hard to reason about precisely because those have different cardinality -and one word ([ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md) +and one word ([ADR 0010](../../02-DECISIONS/0010-delivery.md) is the pipeline half of the same confusion). Split it: @@ -230,7 +230,7 @@ the fact that it runs its own development on them is dogfooding, not architectur ## What agents are, structurally Self-improvement and self-healing are not a tier. Agents are participants -([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)) that hold identity in +([ADR 0003](../../02-DECISIONS/0003-agents-are-persistent-employees.md)) that hold identity in tier 2, act through tier 3 like any other caller, and run as workloads in tier 4. This matters for one reason: **an agent must not have a privileged path**. Anything an agent @@ -241,7 +241,7 @@ no human checkpoint. ## How this is tested -The lab ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)) raises the +The lab ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) raises the tree above on one machine: virtual machines as nodes, a real overlay between them, the real substrate bundle, the real control plane, the real delivery path. @@ -254,7 +254,7 @@ being the one thing nobody exercises until it breaks. - Where the record lives. It is infrastructure by shape and domain by content, and putting it in the substrate risks recreating a circularity in the one place the design just removed one. -- Whether tier 2's contexts are one repository or several. Open from ADR 0015 already. +- Whether tier 2's contexts are one repository or several. Open from ADR 0001 already. - Whether an `edge` node is in the inventory or merely present — which decides whether "node" is one concept or two. - The migration. Nothing here says how today's mesh becomes this, and the skeleton is worth @@ -266,13 +266,13 @@ The tier-0 binary was first called `mesh-agent`, because "node agent" is the ref else in the industry. That is wrong here, and wrong in the specific way [`how-we-build.md`](../../00-META/how-we-build.md) §4 exists to catch: **Agent** is a first-class concept in this mesh — a participant, some of whom are human, holding identity and -memory ([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)). One document +memory ([ADR 0003](../../02-DECISIONS/0003-agents-are-persistent-employees.md)). One document carried both meanings. It is the same failure as the anatomy naming in the current runtime: an evocative domain word pointing at infrastructure. -[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) supplies the fix in +[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) supplies the fix in its own title — *the mesh brokers capabilities; nodes host; agents think.* Three verbs, three components: the control plane **brokers** (`mesh-control`), the tier-0 binary **hosts** (`mesh-host`), the participant **thinks** (`agents`, untouched). diff --git a/01-RESEARCH/007-provisioning-as-the-core/00-overview.md b/01-RESEARCH/007-provisioning-as-the-core/00-overview.md index f331693..aef8c60 100644 --- a/01-RESEARCH/007-provisioning-as-the-core/00-overview.md +++ b/01-RESEARCH/007-provisioning-as-the-core/00-overview.md @@ -3,7 +3,7 @@ status: active initiated: 2026-08-23 touches: - 03-DESIGN/00-as-is/03-provisioning.md - - 02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md + - 02-DECISIONS/0009-modules-and-the-graph.md - 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md became: [] --- @@ -14,7 +14,7 @@ became: [] Provisioning is the mechanism the whole mesh rests on: a module declares what it needs, and the mesh makes it exist, generates the credential, records the grant, and puts the values where the -module will read them. [ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md) +module will read them. [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) calls it the mesh's core concern rather than its plumbing. [Research 006](../006-mesh-from-scratch/code-skeleton.md) then asks it to carry **more**: the diff --git a/01-RESEARCH/008-delivery-coordinator/00-overview.md b/01-RESEARCH/008-delivery-coordinator/00-overview.md index 9bc14a0..21ec3b7 100644 --- a/01-RESEARCH/008-delivery-coordinator/00-overview.md +++ b/01-RESEARCH/008-delivery-coordinator/00-overview.md @@ -1,12 +1,14 @@ --- -status: active +status: graduated initiated: 2026-08-23 touches: - 03-DESIGN/00-as-is/04-delivery.md - - 02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md - - 02-DECISIONS/0013-an-artifact-is-build-output.md + - 02-DECISIONS/0010-delivery.md + - 02-DECISIONS/0010-delivery.md - 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md -became: [] +became: + - 02-DECISIONS/0010-delivery.md + - 02-DECISIONS/0010-delivery.md --- # 008 — The coordinator: a change checked in becomes a deployed state @@ -41,13 +43,59 @@ Research 006 adds a requirement the current design does not have: the coordinato **before the mesh is self-hosting**, when source and artifacts come from outside, and keep working across the transition to self-hosted providers. -## The questions +## What it became + +*Closed 2026-08-28.* All six questions are answered, by two records, and the second exists +because the first was honest about what it did not fix. + +**Does the coordinator dispatch stages, or converge nodes on a declaration?** — *Converge.* +[ADR 0010](../../02-DECISIONS/0010-delivery.md): a pipeline ends when the +declaration is updated, and the host applies it and reads back — so the reporter is the applier. + +**Does the three-silo split survive?** — *Yes, with the third redefined.* The cardinality +observation holds; the third silo is not a stage any more. + +**How does a change become a pipeline, reliably?** — *It does not become a pipeline at all.* +[ADR 0010](../../02-DECISIONS/0010-delivery.md) applies 0058's +move one level up: the control plane holds what source exists and what has been built, and builds +the difference. **An event makes it fast; nothing makes it necessary.** The failures this effort +catalogued — a truncated commit list, a broken path match — become latency rather than silence. + +**What is a deployed state?** — *Two comparisons, not an event.* Does every node's reported state +match what is declared, and is what is declared built from current source? A milestone can be +claimed by something that did not check; a comparison cannot. + +**What produces a verdict, and what is it about?** — *An artifact, and it gates eligibility.* The +mesh must not converge onto something broken, so an artifact may be declared only once the lab +has judged it fit. Sharper than the question expected: a verdict is a property an artifact has, +not a report about a run. + +**How does delivery work before self-hosting?** — *It mostly stops being a question.* A reconciler +reads source and writes artifacts; where those live is a binding, external first and internal +later. The transition looked hard because a pipeline's stages name their targets. + +## What this effort was right about + +Its first question — *what is a deployed state, and how does the mesh know it is in one* — was +marked "everything follows from this", and everything did. Both records above are answers to it: +0058 makes the applier the reporter, and 0063 makes currency a comparison. The effort put the +load-bearing question first. + +## What is NOT closed by this + +[ADR 0010](../../02-DECISIONS/0010-delivery.md) names four costs +and one of them is a real risk rather than a trade: **a reconciler that cannot reach its target +retries forever, and without something that notices, the failure is silence** — which is the +fault this effort exists to catalogue, reintroduced in a new place. That belongs to observability +and it is not designed. + +## The questions (all answered above) | Question | Why it matters | |---|---| | What is a **deployed state**, and how does the mesh know it is in one? | Everything follows from this. If a stage reports transport, "deployed" is a claim nobody checked. A desired-state model with reconciliation gives a different answer from a job-completion model. | | Does the coordinator dispatch **stages**, or converge nodes on a **declaration**? | The current model is a state machine over stages. The alternative is that a node is told what should be true and reports what is. The second makes drift visible; the first cannot see it. | | How does a change **become** a pipeline, reliably? | Detection has failed for reasons unrelated to the change, silently. | -| What produces a **verdict**, and what is it a verdict about? | Ties to the lab ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)) and to a module carrying its own assertions. | +| What produces a **verdict**, and what is it a verdict about? | Ties to the lab ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) and to a module carrying its own assertions. | | How does delivery work **before self-hosting**, and across the transition? | From research 006: source and artifacts start external and are re-bound to internal providers. The coordinator has to be indifferent to which. | -| Does the **three-silo** split survive the artifact/part split? | [ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md) is cardinality-driven, and research 006 renames the thing the cardinality is about. | +| Does the **three-silo** split survive the artifact/part split? | [ADR 0010](../../02-DECISIONS/0010-delivery.md) is cardinality-driven, and research 006 renames the thing the cardinality is about. | diff --git a/01-RESEARCH/009-migration/00-overview.md b/01-RESEARCH/009-migration/00-overview.md index 832d76a..e87fe36 100644 --- a/01-RESEARCH/009-migration/00-overview.md +++ b/01-RESEARCH/009-migration/00-overview.md @@ -4,7 +4,7 @@ initiated: 2026-08-23 touches: - 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md - 03-DESIGN/01-to-be/01-end-to-end-testing.md - - 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md + - 02-DECISIONS/0016-the-lab.md - 03-DESIGN/00-as-is/00-overview.md became: [] --- @@ -28,7 +28,7 @@ Recorded because incremental is the reflex answer and it is wrong in this case. requirements — none of these can half-apply. Running both models at once means the old one's assumptions keep constraining the new one, which is how a migration becomes permanent. - **Nothing external depends on it.** No users outside the operator, no service level to hold. -- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)). +- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the same risk as one performed for the first time on the real mesh. This is also why the lab is phase 0 rather than a verification step later: the new mesh is *developed* inside it, so by @@ -70,7 +70,7 @@ than a discovery. | Phase | What | Done when | |---|---|---| -| **0** | **Build the lab's bootstrap scenario** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. | +| **0** | **Build the lab's bootstrap scenario** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. | | A | Build tier 0, **inside the lab**. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. | | B | Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. | | C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. | diff --git a/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md index db40f90..44210d9 100644 --- a/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md +++ b/01-RESEARCH/010-lab-inner-loop-cost/00-overview.md @@ -3,8 +3,8 @@ status: active initiated: 2026-08-24 touches: - 03-DESIGN/01-to-be/03-scenario-lifecycle.md - - 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md - - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md became: [] --- @@ -23,7 +23,7 @@ Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measuremen ## Why it matters -[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the +[ADR 0016](../../02-DECISIONS/0016-the-lab.md) makes the bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that raising a node from nothing stops being the least-exercised path and becomes the most-exercised one. **That argument is only true if raising and resetting are cheap.** A loop that costs diff --git a/01-RESEARCH/010-lab-inner-loop-cost/measurements.md b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md index a690054..ee3d9a8 100644 --- a/01-RESEARCH/010-lab-inner-loop-cost/measurements.md +++ b/01-RESEARCH/010-lab-inner-loop-cost/measurements.md @@ -13,7 +13,7 @@ image. | Fact | Value | Consequence | |---|---|---| -| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty | +| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-the-lab.md) is not paying an emulation penalty | | Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot | | Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on | | btrfs kernel module | **available** | the kernel can do it | @@ -72,7 +72,7 @@ worst, before any of the mesh's own work begins. **This is too slow for an inner loop**, and the reason is not the design. -[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that +[ADR 0016](../../02-DECISIONS/0016-the-lab.md) argues that making the bootstrap path the inner development loop turns the least-exercised code in the system into the most-exercised. That argument holds only while resetting is cheap. At a minute and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around @@ -120,7 +120,7 @@ A four-machine reset-and-rerun cycle, the operation the inner loop repeats most: | **cycle** | **~90 s, unbounded at worst** | **~15 s, dominated by boot** | At fifteen seconds, dominated by a boot that cannot be avoided, the inner loop is viable and -[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)'s argument holds. +[ADR 0016](../../02-DECISIONS/0016-the-lab.md)'s argument holds. At ninety it did not. ### One honest counter-observation diff --git a/01-RESEARCH/011-the-module-graph/00-overview.md b/01-RESEARCH/011-the-module-graph/00-overview.md index 2fef368..b9a9ff0 100644 --- a/01-RESEARCH/011-the-module-graph/00-overview.md +++ b/01-RESEARCH/011-the-module-graph/00-overview.md @@ -1,10 +1,15 @@ --- -status: active +status: graduated initiated: 2026-08-25 +became: + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0008-a-context-owns-its-store.md + - 03-DESIGN/01-to-be/06-the-control-plane.md + - 03-DESIGN/01-to-be/07-the-substrate.md touches: - - 02-DECISIONS/0002-everything-is-a-module.md - - 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md - - 02-DECISIONS/0036-a-node-is-a-managed-machine.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0004-a-node-and-how-it-joins.md - 03-DESIGN/00-as-is/02-modules-and-manifests.md - 03-DESIGN/00-as-is/10-module-catalogue.md - 04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md @@ -12,15 +17,69 @@ touches: # 011 — The module graph +> **A third edge was found after this graduated.** This effort established *presence* and +> *instantiation*, and both are **runtime** edges — they answer *what does this need in order to +> run*. Delivery needs a different question answered — *what has to be rebuilt when this changes* +> — and that is a **build** edge, fixed inside an artifact rather than negotiated when it runs. +> Recorded by [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md), which also +> notes what this effort's three entities turn out to be good for +> ([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)). + ## What is being investigated Whether the catalogue's missing structure is a **graph** — modules declaring what they need, what they offer, and what they exclude — and what that replaces. -[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) proposes +**[`worked-provider.md`](worked-provider.md) works one module through completely**, and breaks +the tidy version. A database is nine things, not one — and a small game asking the mesh for its +own database shows there are **two kinds of edge**: *presence*, where the thing must exist, and +*instantiation*, where a provider makes something for a consumer and hands back credentials. +Instantiation implies presence and not the reverse. The current system already had this split +and the design had collapsed it. + +**[`features.md`](features.md) answers what happens to `feature`.** It is one word for four +things spanning three tiers — artifacts built once per version, resources applied to a machine, +actions run against something that is not this machine, and checks that are requirements in +disguise. Measured: every one of the twenty-one handlers implements all six stages, so `configs` +has a build stage with nothing to build and `npm` has a start stage with nothing to start. That +emptiness is the conflation, and it is why a stage that did nothing and a stage that failed look +alike. Nothing replaces it, because it was never one concept. **One property is worth keeping: +content is detected, relationships are declared.** + +**[`cases.md`](cases.md) enumerates what a module can be** — twenty kinds of thing the mesh has +to install, run, own or know about — and extracts the axes a manifest must express. Two of those +axes appear in no current thinking: **how many instances** a thing may have, and **whether two +can coexist**. + +**The design is in [`proposal.md`](proposal.md): one kind of edge.** A module provides names +and requires names, and that single relation absorbs requiring a module, requiring a resource, +and the interface-and-adapter idea. An abstract name is legitimate **only where providers are +genuinely substitutable** — `terminal` passes, `database` does not, because a consumer speaking +Postgres does not speak MongoDB. The adapter is what creates an interface; without one there is +a **tag**, which describes and does not bind. +A **node provides names too**, which makes capability checking stop being a separate mechanism +and makes the host's own capability report an input to resolution rather than something a +person reads. + +[`analysis.md`](analysis.md) measured the current catalogue. Its value to the design is two +lessons rather than its machinery — a field that means *depends on* should say so, and +placement does not belong in a manifest — and the rest is recorded as as-is evidence. + +**Measured, and the premise was wrong for the existing system: the graph is not missing there.** +[`analysis.md`](analysis.md) — 126 manifests, 103 edges, no cycles, nothing dangling, and a +resolver that topologically sorts them, already called by the tool loader, the installer and +the delivery coordinator. What the effort assumed would need building is a thing to call. + +What survives is narrower: three declarations that do not exist (`excludes`, a required node +capability, an interface with adapters), and two defects worth fixing whatever else is +concluded — `provider:` is a dependency edge that is not read as one, which makes the closure +for a working mesh come out without a database; and the resolver continues past a cycle and +past a missing dependency, contrary to ADR 0001. + +[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) proposes grouping modules by domain. [Research 005](../005-domain-grouping/analysis.md) measured that proposal and found its evidence holds in exactly one place — reachability — which -[ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) has since absorbed +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) has since absorbed into the host. The measured case for domain grouping has therefore been consumed by a decision taken for unrelated reasons, and what remains is fifty modules that co-change with nothing. @@ -43,7 +102,7 @@ and abandoned in favour of one concept with facets, for a reason worth keeping: Filing decisions that follow from nothing are the disease research 005 measured. A second taxonomy would reproduce it. -So [ADR 0002](../../02-DECISIONS/0002-everything-is-a-module.md) survives, and the question +So [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) survives, and the question becomes what a module must be able to **declare**. ## The shape being investigated @@ -52,7 +111,7 @@ Five declarations, of which two exist today. | Declaration | Today | Notes | |---|---|---| -| **requires a resource** — a database, a bucket | yes | provisioning, [ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md) | +| **requires a resource** — a database, a bucket | yes | provisioning, [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) | | **provides a resource** | yes | as above | | **requires another module** | **no** | the dependency edge — the graph's substance | | **excludes another module** | **no** | installing A makes B unavailable | @@ -106,12 +165,39 @@ integration being wrong looks like from the outside. ## Open questions +Struck-through rows are answered, with where. The rest are live. + +### Settled + +| Question | Answer | +|---|---| +| ~~What does the graph **delete**?~~ | For the existing system: nothing, it is already there ([`analysis.md`](analysis.md)). For the design: the module/resource distinction, the interface as a kind of thing, capability checking as a separate mechanism, domain grouping, and — the first clear deletion — **grant kinds**, once a module may only be granted what it exclusively owns ([`worked-provider.md`](worked-provider.md)). | +| ~~Is an interface a module, or a name?~~ | A **name**, and only where providers are genuinely substitutable. The adapter is what creates one; without an adapter there is a **tag**, which describes and does not bind ([`proposal.md`](proposal.md)). | +| ~~Where do domain modules fit?~~ | They do not. There is core infrastructure — concrete modules named individually, not flavourable, nothing standing in front of them. | +| ~~What happens to domain grouping?~~ | Superseded. Folders assert relationships; edges record them. What grouping was for is a tag and a query. [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) is `proposed` and should be superseded rather than narrowed. | +| ~~Is there one kind of edge?~~ | **No — two.** *Presence*, where a thing must exist, and *instantiation*, where a provider makes something for a consumer and hands back credentials. Instantiation implies presence, not the reverse. | +| ~~When two modules provide one name, who chooses?~~ | Neither the consumer naming a node nor the consumer not caring. The consumer declares the **scope of its own need** — shared across its instances, or one each — the mesh binds, and the binding is written down and sticky. Where it is written follows the scope. | +| ~~Can several modules share one database?~~ | **No.** A module is granted only what it exclusively owns — no shared writes and no read role on another's store, because reading couples you to its layout just as firmly. | +| ~~What about a dashboard reading a dozen stores?~~ | **The rule is about contexts, not processes.** The mesh's own board reading the mesh's own store is the mesh showing its own data — not a boundary crossing. Everything inside a context reads its store freely; what is forbidden is a *different* context reading it. An earlier answer here was wrong. | +| ~~Can every registry consumer be served another way?~~ | **Largely dissolves.** Of eighteen direct consumers, the owner keeps its database, node appliers are already stopped by ADR 0005, and the bulk are **foreign tenants** — thirteen tables across three contexts — who need to move out rather than read differently. | + +### Live + | Question | Why it is open | |---|---| -| What does the graph **delete**? | If modules gain declarations and lose nothing, this is motion rather than progress. The effort has not finished until it names what stops existing. | -| Where does resolution happen — mesh or platform package manager? | The mesh must model mesh-level edges. Whether it also resolves operating-system packages, or delegates, decides whether a solver has to be written. | -| Is an interface a module, or a name? | Arch makes it a name that packages claim. Making it a module gives it a manifest, an owner and a place to document the contract — and a thing with no implementation to install. | -| What does an exclusion mean for something already installed? | Refuse the install, or make the conflict visible and let it be decided. The second is a policy surface; the first is a package manager. | -| Does node adoption scan for capabilities, applications, or both? | The operator proposes scanning an adopted node and enabling what it finds. Under the working position above, the scan is for capabilities — but a machine with a terminal already installed is also a module already satisfied, and whether that is adoption or drift is undecided. | -| One installation image, or several? | Proposed: pre-built images carrying different capability sets, so a machine is adopted quickly. Several images bake capability sets at image time, which is the filing problem in a new form and reintroduces what detection exists to avoid. One image carrying the host and nothing else is [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md)'s *one binary installed by hand*, automated. The effort should settle which. | -| What happens to domain grouping? | [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) is still `proposed`. If the graph is the answer, 0017 is superseded rather than narrowed — its text is never edited. | +| How many instances should a module have? | Not derivable and not global: one-per-mesh is right for the broker and wrong for a store a disconnected node needs. It is per-module and nothing in the schema says it. | +| Can two of something coexist? | `excludes` covers part of it. Two terminals are fine, two things wanting one port are not, two brokers might be either. | +| What does a provider hand back? | Credentials and an address for a store; a command for a terminal. Same relation, different shape crossing it. | +| Is provisioning one mechanism or two? | The mesh's own registry is provisioned **before there is a mesh**, so provisioning is part of the bootstrap and part of what the carried bundle expresses. At bootstrap the store is local; afterwards it is on another node. Same operation, both sides of a tier boundary. | +| Is a tool surface one relation with two audiences, or two? | 56 of 126 modules carry tools — more than carry a service — and what consumes them is an **agent**, not a module. | +| ~~Do the remaining cross-context reads want an interface or events?~~ | **Derived, not chosen.** Neither is SQL — that only ever runs against your own store. ADR 0004 makes disconnection ordinary, so anything that must work while disconnected cannot use a request and needs a local copy: a subscription. Anything where a stale answer is worse than none cannot use a subscription. | +| What does a consumer do about events it missed while disconnected? | Replay from a point, ask once for a full picture and resume, or rebuild. The question every projection has, and unanswered here. | +| What happens to a grant when its consumer is removed? | Dropping is data loss; keeping is a leak. ADR 0005's removal rule does not obviously carry, because the thing lives inside another module's state. | +| Is a declaration composed per node, from what that node reported? | Some configuration follows the hardware. Either the host fills a blank — deciding, against ADR 0005 — or the control plane composes from the node's inventory first. | +| Would `excludes` and capability requirements actually be used? | Zero manifests declare either, which is equally consistent with *nobody needs them* and *nobody can express them*. | +| What does an exclusion mean for something already installed? | Refuse the install, or surface the conflict and let it be decided. | +| Are tiers a view of the graph, or a constraint on it? | If a tier is a computed level the word is a convenience. If *a tier may depend only on tiers below it* is to be enforced, it is a constraint and must be stated as one. | +| Where does resolution happen — the mesh, or the platform's package manager? | The mesh must model mesh-level edges. Whether it also resolves operating-system packages decides whether a solver gets written. | +| How far may the control plane be split? | A single board over several contexts works because they are contexts *inside* one control plane with one interface. If a context becomes its own deployable with its own interface, the board is coupled to N of them and the composition has nowhere to live that tier 3 permits. A constraint on splitting, worth knowing before splitting. | +| What does the pipeline schedule, once features are gone? | It schedules features today. The four categories they split into have different lifecycles, so the unit of work differs for each and needs naming. | +| Should placement leave the catalogue? | A provision pins itself to a named node in the manifest. Placement is an inventory decision, and having it in the catalogue means a second node cannot provide the mesh's store without editing its consumer. | diff --git a/01-RESEARCH/011-the-module-graph/analysis.md b/01-RESEARCH/011-the-module-graph/analysis.md new file mode 100644 index 0000000..ef42317 --- /dev/null +++ b/01-RESEARCH/011-the-module-graph/analysis.md @@ -0,0 +1,138 @@ +# The graph is not missing + +Measured against `origin/main` of the code repository, 2026-08-26. Every manifest, read through +git refs rather than a checkout. + +The effort was opened to ask whether the catalogue's missing structure is a graph. It is not +missing. **It exists, it is healthy, and three separate parts of the system already use it.** + +That is the finding, and it changes what is worth asking. + +## Finding 1 — the graph is already declared, and it is clean + +| | | +|---|---| +| manifests | 126 | +| declare a dependency on another module | 64 | +| declare a requirement on a provision | 11 | +| **edges** | **103** | +| declare neither | 62 — 49% | +| **cycles** | **0** | +| **dependencies declared but absent** | **0** | +| deepest chain | 5 | + +Half the catalogue is unconnected, which matches +[research 005](../005-domain-grouping/analysis.md)'s finding that fifty modules co-change with +nothing. The connected half is well formed: no cycles, nothing dangling. + +## Finding 2 — it is already resolved, and already used + +The platform SDK carries a dependency resolver that topologically sorts modules, and it is +called from three places: the tool loader at startup, the installer when syncing modules onto a +node, and the delivery coordinator when expanding what a change affects. + +It also already does something the effort assumed would need designing: **a requirement on +another module's provision is treated as an implicit edge to that module**, so a consumer does +not have to declare the same relationship twice. + +So *ordering by the graph* — which +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) says +the control plane will do — is not a thing to build. It is a thing to call. + +## Finding 3 — the most important edges in the mesh are invisible + +The one place the graph is wrong, and it is wrong about the substrate. + +A module that needs a database declares it like this: + +```yaml +provisions: + - name: mesh-db + provider: postgres # ← a dependency on the postgres module + node: # ← and where it must run +``` + +`provider:` names a module. It is a dependency, declared, in the manifest — and it sits inside +`provisions:`, which is what a module *offers*. The resolver reads `dependencies:` and +`requires:`, so it never sees it. + +| | | +|---|---| +| provider references that name a real module | 4 | +| **invisible to the resolver** | **3, across 2 modules** | + +Three edges is nothing, and they are the mesh's own database, the mesh's own broker, and the +work engine's database. The most load-bearing relationships in the system are the ones the +graph cannot see. + +**The consequence is measurable.** Computing what a working mesh needs, from the declared +graph: + +``` + registry → sdk → mesh → meshware 4 modules, 4 levels +``` + +No database. No broker. A closure that is arithmetically correct and obviously wrong, and wrong +for exactly one reason: a field that means *depends on* is not read as one. + +This is [04-ISSUES/003](../../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) +again, in a new place — not a key nothing reads, but a key read as something other than what it +means. + +## Finding 4 — the resolver continues past faults it should stop on + +Two behaviours, both contrary to +[ADR 0010](../../02-DECISIONS/0010-delivery.md): + +- **A cycle warns and falls back to input order.** A cycle means no correct order exists; the + resolver proceeds with an arbitrary one and logs a line. +- **A dependency that does not exist warns and continues.** The validation is documented as + *non-fatal, logged as warnings*. + +Neither has fired in the current catalogue — there are no cycles and nothing dangling — which +is why nobody has noticed. They are latent, and they are in the component that +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) makes +responsible for the ordering a host will apply without question. + +## Finding 5 — placement is decided in the catalogue + +`node:` in a provision pins it to a named node, in the manifest. Two modules do this today, +and they are the substrate ones. + +Which node runs what is an inventory and placement decision — tier 2 by the skeleton's own +test. Having it in a manifest means the catalogue decides placement, and a second node cannot +provide the mesh's database without editing the module that consumes it. + +## What this means for the effort + +**The opening question — "what does the graph delete?" — has an answer: nothing, because the +graph is already there.** The premise was wrong, and finding that out is the effort's first +result rather than a setback. + +The questions that survive are narrower and answerable: + +| Missing declaration | Manifests using it today | +|---|---| +| `excludes` — installing A makes B unavailable | **0** | +| a required node capability | **0** | +| an interface, with adapters providing it | **0** | + +Those are what a graph would *add*. What it would delete is a different and smaller list, and +the honest version of it is: nothing yet. + +**And two defects worth fixing regardless of what else this effort concludes:** + +1. `provider:` is a dependency edge and is not read as one. Fixing it makes the closure correct + — which is what [research 012](../012-the-minimum-viable-node/00-overview.md) needs in order + to answer what a one-node mesh requires. +2. The resolver continues past a cycle and past a missing dependency. Both should refuse. + +## What was not measured + +- **Whether the missing declarations would be used.** Zero manifests declare exclusions or + capabilities, and that is equally consistent with *nobody needs them* and *nobody can express + them*. Nothing here separates those. +- **Whether the interface-and-adapter idea has a consumer.** It is a good shape, and it is + argued for rather than measured. +- **What the closure should be.** Finding 3 says the computed one is wrong. It does not say + what the right one is; that needs the fix first. diff --git a/01-RESEARCH/011-the-module-graph/cases.md b/01-RESEARCH/011-the-module-graph/cases.md new file mode 100644 index 0000000..d4e7c73 --- /dev/null +++ b/01-RESEARCH/011-the-module-graph/cases.md @@ -0,0 +1,128 @@ +# What a module can be + +Every kind of thing the mesh has to install, run, own or know about, before deciding what a +manifest says. Written to be argued with: a case here that turns out not to exist should be +struck, and one that is missing is a hole in whatever schema follows. + +The hard cases are at the end, and they are the point. + +## The ordinary cases + +**1 — A supervised service.** A container the mesh runs and keeps running. Nobody starts it; +it is simply up. *A relational store, a message broker, an object store, a dashboard.* + +**2 — A system package with configuration.** Not a container. Installed into the machine, +configured through files, run by the service manager. *A firewall, a resolver, an overlay.* +Note: [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) says applying +these is the host's job — so what the module contributes is the *deciding*, not the doing. + +**3 — An application a person launches.** Installed on a node, started by a human, running only +while they use it. *An editor, a chat client, a file manager, a terminal.* + +**4 — A command-line tool.** Installed, on the path, run when invoked. No service, no window. +*A formatter, a query client, a backup utility.* + +**5 — A library.** Never runs at all. Consumed at build time by other modules. *An SDK.* + +**6 — A one-shot task.** Runs once, changes something, exits. *A schema migration, a data +import, a seed.* + +**7 — A scheduled task.** Runs repeatedly on a timer, exits each time. *A backup, a prune, a +report.* + +**8 — An adapter.** Exists to make several unlike things look alike behind one name. *A model +provider behind an assistant interface.* + +**9 — A standalone application in its own repository.** Same shape as any of the above; the +difference is only where its source lives +([ADR 0015](../../02-DECISIONS/0015-applications-live-in-their-own-repository.md)). Worth +listing because a schema that assumes a monorepo path would exclude it. + +## The cases that break a naive schema + +**10 — Something that is a service *and* an application.** A git forge is consumed by other +modules as a remote and a registry, *and* operated by a person through a web interface. A web +analytics service grants a tracking identity *and* is a dashboard somebody reads. Neither is a +service-or-application choice; both are true simultaneously. + +**11 — Something that provides to others *and* consumes from others.** The store provides +databases and needs a filesystem. The forge provides a registry and needs a database. Provider +and consumer are not kinds of module; they are ends of edges. + +**12 — Something the mesh installs that then becomes a node capability.** The container runtime +is installed *by* the mesh, and once it works, the node **provides** `container-runtime` to +everything else. So a module can change what its node provides. The node's provides-list is +therefore partly derived from what is installed on it, not only detected from what was already +there — and the two have to agree. + +**13 — Something that must be adopted rather than installed.** The machine already has the +package manager, the container runtime, possibly the version control system, each with +configuration somebody chose. The module does not install it; it takes it over +([research 012](../012-the-minimum-viable-node/00-overview.md)). + +**14 — Something that is a set, not a thing.** *Core infrastructure* is not installable — it is +a name for the concrete modules a working node needs. Whether that is a module whose only +content is `requires`, or a query over the graph, or a pinned list outside the catalogue, is +undecided and is one of the sharper questions here. + +**15 — Something with one instance for the whole mesh.** There is one mesh database, not one per +node. Assigning it to two nodes is not redundancy, it is two meshes. Contrast with a terminal, +where per-node is the only sensible reading. + +**16 — Something that may be installed several times over.** Two terminals coexist happily. Two +things wanting port 443 do not. Two message brokers might be fine or might be a split brain, +and nothing in *provides* and *requires* distinguishes those. + +**17 — Something that is not software at all.** A firewall policy. A DNS record. A certificate. +It owns no binary, runs nothing, and is entirely *desired state* — which is the one case that +fits the host's declaration model exactly and fits an installable-package model not at all. + +**18 — An agent.** The mesh's own premise is that agents are participants. An agent has an +identity, a licence, a node it runs on, and work it does. Whether that is a module, a record in +the control plane, or something else is not obvious, and getting it wrong shapes everything +about how agents are assigned. + +**19 — The host itself.** Tier 0 installs everything else and is installed by hand. It is not a +module, and a schema that cannot say so has a bootstrap problem hiding in it. + +**20 — Something the mesh depends on and does not control.** A domain registrar, an upstream +resolver, a certificate authority, an electricity supply. Almost certainly not modules — but +the mesh's health depends on them, and they are the reason a node can be perfectly configured +and still not work. + +## What varies across the cases + +The list above matters less than this. These are the axes a manifest has to express, and each +one is a question the schema must answer or deliberately refuse. + +| Axis | Range | Sharpest case | +|---|---|---| +| **Does it run?** | supervised · launched by a person · once · on a timer · never | 5, 6, 7, 17 | +| **Who starts it?** | the mesh · a human · nothing | 1 vs 3 | +| **Where does it come from?** | container image · system package · our source · already on the machine | 12, 13 | +| **Does it provide to other modules?** | a resource · an abstract name · nothing | 8, 11 | +| **Does it hold state?** | yes, and it matters where · no | 1 vs 3 | +| **How many instances?** | one per mesh · one per node · many per node | 15, 16 | +| **Can two coexist?** | yes · no · only with different settings | 16 | +| **Is it installed or adopted?** | installed · adopted · either, depending on the machine | 13 | +| **Does installing it change what the node provides?** | yes · no | 12 | + +**Two of these are not in any current thinking**, and both come from the hard cases: +*how many instances* (15) and *can two coexist* (16). `excludes` covers part of the second and +nothing covers the first. + +## Questions the cases raise + +- **Is "runs" a property or a kind?** The axes suggest a property — one schema, with a field + saying how it runs, `never` included. The alternative is several kinds of module with + different schemas, which is the taxonomy [research 011](00-overview.md) already rejected once + for services and applications. +- **Is an agent a module?** (18) If yes, the schema carries identity and licensing. If no, the + mesh has two catalogues. +- **Is core infrastructure a module?** (14) A module whose only content is `requires` is either + elegant or a category pretending to be a thing — the same trap `database` was. +- **What names one-per-mesh?** (15) Nothing in provides, requires or excludes says it, and + getting it wrong means two of something that must be one. +- **Where does a policy live?** (17) Pure desired state fits the host's declaration exactly. + Whether it is a module at all, or something the control plane derives and no catalogue entry + exists for, is open. diff --git a/01-RESEARCH/011-the-module-graph/features.md b/01-RESEARCH/011-the-module-graph/features.md new file mode 100644 index 0000000..84c1216 --- /dev/null +++ b/01-RESEARCH/011-the-module-graph/features.md @@ -0,0 +1,131 @@ +# What a feature is, and what it splits into + +Measured against `origin/main` of the code repository, 2026-08-26. + +The operator wants the feature concept gone. +[Research 006](../006-mesh-from-scratch/00-overview.md) left it open — *"does `feature` +survive? The skeleton splits it in two and argues the conflation is what makes the delivery +pipeline hard to reason about."* + +This is what it actually is, and what it turns into. + +## What it is today + +A **feature** is a kind of content a module can carry, detected from what its directory +contains rather than declared. Twenty-one of them, each with a handler that owns its whole +lifecycle: + +``` +configs dist-assets events hooks migrations migration-artifact npm npm-install +prerequisite-env prerequisite-packages prerequisite-provision provision-migrations +provision-seeds seeds service systemd tools verifiers verify vhost detect +``` + +Each moves through **six stages**: build, publish, install, configure, start, verify. + +## Finding — every handler implements every stage + +The structural evidence, and it is not a style problem. + +| handler | stages it implements | +|---|---| +| `configs` | build · configure · install · start · verify | +| `service` | build · configure · install · start · verify | +| `systemd` | build · configure · install · start · verify | +| `tools` | build · configure · install · start · verify | +| `vhost` | build · configure · install · start · verify | +| `migrations` | build · configure · install · start · verify | +| `npm` | build · publish · install · configure · start | + +`configs` writes files onto a node. It has nothing to build, and it has a build stage. +`npm` publishes a package to a registry. It has nothing to start, and it has a start stage. + +**One interface spans build-time and apply-time, so every kind of content must implement both +halves and most of them do nothing in one.** That is the conflation, and the emptiness is what +makes the pipeline hard to reason about: a stage that does nothing and a stage that failed to +do anything look identical from outside. + +## What it splits into + +The twenty-one are not one kind of thing. They are four, and they belong to different tiers. + +**Artifacts — built once per version, then published.** `npm`, `dist-assets`, +`migration-artifact`. Nothing about a node is involved; the output is a thing that exists in a +registry. **Tier 2, delivery.** + +**Resources — desired state on a machine.** `configs`, `service`, `systemd`, `vhost`, `tools`. +Applied, converged, idempotent — which is exactly what +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) +already describes and what the host already does. **Tier 0.** + +**Actions — run once, against something that is not this machine.** `migrations`, `seeds`, +`provision-migrations`, `provision-seeds`, `hooks`. A migration runs against a database, and the +database may be on another node entirely. Neither an artifact nor node state, which is why they +sit awkwardly in a scheme built for both. **Tier 2, and the operator wants seeds gone.** + +**Checks — assertions, not changes.** `prerequisite-env`, `prerequisite-packages`, +`prerequisite-provision`, `verify`, `verifiers`, `detect`. The prerequisites are **requirements +in disguise** — a module saying what must be true before it can be installed, which is precisely +what an edge in the graph says. The verifiers are the read-back the host already performs. + +## So the answer is: it splits, and the split is a tier boundary + +**`feature` is one word for four things spanning three tiers.** That is why every handler +implements every stage, why half of them are empty, and why the pipeline is hard to reason +about. + +Nothing replaces it, because it was never one concept: + +| Was a feature | Becomes | Whose | +|---|---|---| +| npm, dist-assets, migration-artifact | an **artifact** | delivery | +| configs, service, systemd, vhost, tools | a **resource** in a declaration | the host | +| migrations, hooks | an **action** against something else | delivery | +| seeds | *nothing* — the operator wants them gone | — | +| prerequisite-* | an **edge** in the graph | the catalogue | +| verify, verifiers | the host's **read-back** | the host | + +## Tools are a facet, and their consumer is not a module + +The most common content in the catalogue: 56 of 126 modules carry a tool surface, more than +carry a service. + +It survives the split, and it does not fit either half cleanly. A tool is not an artifact and +not node state — it is a **contract the mesh publishes on a module's behalf**, and what consumes +it is an **agent**, not another module. + +That is a second audience, and the design has only described one. A module `provides` things +other modules require; a module also `provides` things agents call. Same word, different +consumer, different lifecycle — a tool appears when the module is assigned somewhere and +disappears when it is not, and nothing in the graph edges says so. + +Whether that is one relation with two audiences or two relations is undecided, and it is the +kind of question that is cheap now and expensive later. + +## What is lost, and should not be + +**Detection.** Features are detected from the directory rather than declared, and the reason is +good: *a declared list and the directory it describes drift, and the directory is the one that +is true.* That property is worth keeping whatever the concept is called — a module that says it +has migrations and has none, or has them and does not say so, is a fault nobody sees until it +matters. + +Under the split, detection still applies: what a module **contains** is read from what is there. +What it **requires**, **provides** and **excludes** is declared, because none of that is visible +in a directory. + +That line is worth stating precisely, because it is the one the current design got right: +**content is detected, relationships are declared.** + +## Open + +- **Is an action a resource?** A migration is not node state and not an artifact. It could be a + resource type the host applies with the target being a database rather than a machine — which + would collapse the third category into the second, at the cost of the host reaching something + that is not the machine it is on. That cost looks too high, but it has not been argued. +- **Where do hooks go?** They are arbitrary code a module runs at a stage — the escape hatch. A + design with no escape hatch is either very good or has not met reality yet, and this one has + not. +- **What does the pipeline schedule, once features are gone?** It currently schedules features. + If the four categories have different lifecycles, the unit of work is different for each, and + what the coordinator orchestrates needs naming. diff --git a/01-RESEARCH/011-the-module-graph/proposal.md b/01-RESEARCH/011-the-module-graph/proposal.md new file mode 100644 index 0000000..21e5297 --- /dev/null +++ b/01-RESEARCH/011-the-module-graph/proposal.md @@ -0,0 +1,167 @@ +# One kind of edge + +> **Superseded in part by [`worked-provider.md`](worked-provider.md).** Working postgres through +> completely shows there are **two** kinds of edge, not one: *presence* — the thing exists and is +> reachable, nothing created — and *instantiation* — the provider is asked to make something for +> this consumer and hands back credentials. Instantiation implies presence; presence does not +> imply instantiation. Everything else below stands; the claim in the title does not. + +A design, not an account of what exists. [`analysis.md`](analysis.md) measured the current +catalogue and its value here is two lessons rather than its machinery: a field that means +*depends on* should say so, and placement does not belong in a manifest. + +## The idea + +**A module provides names. A module requires names. That is the only edge.** + +Everything the effort listed as separate declarations turns out to be one relation with +different kinds of name on either end. + +``` +postgres provides postgres +kitty provides kitty, terminal +xterm provides xterm, terminal +anthropic provides anthropic, ai-assistant + +meshboard requires postgres, lavinmq +vscode requires terminal, display-server +``` + +A name is either **concrete** — a module's own name, so `requires: postgres` means that module +and no other — or **abstract**, so `requires: terminal` means whatever provides it. + +That single relation absorbs three things the effort had listed separately: + +| Was going to be | Is | +|---|---| +| requires another module | requires a concrete name | +| requires a resource | requires the module that provides it, by name | +| an interface, with adapters providing it | an abstract name, legitimate only where an adapter makes providers substitutable | + +The interface stops being a kind of module. It is a name with more than one provider *and a +contract they all satisfy*, and nothing has to declare that it is one. + +### An abstract name is only legitimate when providers are actually substitutable + +The test, and it is a strict one: **can a consumer be switched from one provider to another +without changing?** + +`terminal` passes. Anything that runs a command in a terminal works, and a consumer never +learns which one it got. + +**`database` fails, and it is the example worth keeping.** A module connecting to Postgres does +not connect to MongoDB, or to SQL Server, or to MariaDB. Different wire protocol, different +dialect, different driver. A consumer that declared `requires: database` and was handed any of +them would break — so the name promises something no provider can deliver, and the resolver +would satisfy a requirement that is not satisfied. + +That failure has a shape this repository already knows: something declared, accepted, and not +true. It is worse here than usual because the graph would report success. + +**The adapter is what creates an interface.** `ai-assistant` is a legitimate abstract name +exactly when adapters exist to normalise the providers behind it. Without an adapter there is +no interface — there is a category, and a category is not an edge. + +So `postgres` is required by name, and a module that could genuinely work with several stores +requires whichever it actually speaks to. + +### Categories are catalogue metadata, not structure + +*Database* is still a useful word — for finding things, for a person browsing what the mesh can +host, for grouping in an interface. It is a **tag**. + +Tags describe. Edges bind. Keeping them apart is what stops the catalogue acquiring a second +kind of relationship that looks like a dependency and is not — which is what a folder named +after a domain already was. + +## The node provides names too + +The move that makes capabilities stop being a separate system. + +A node's profile is a set of provided names. `display-server`. `container-runtime`. +`amd64`. A module requiring `display-server` is satisfied by **the node**, exactly as a module +requiring `postgres` is satisfied by another module. + +So there is one resolution rather than two: *is this name provided by anything available here?* +A graphical application cannot be installed on a node with no display server for the same +reason, and through the same code, that it cannot be installed without its libraries. + +And it means the host's capability detection — which already reports what a machine can be +asked to do, with a reason for each verdict — **is the node's provides-list**. It was built to +be read by a person; it turns out to be an input to resolution. + +## What a module declares + +```yaml +name: meshboard + +provides: [meshboard] +requires: [postgres, lavinmq, container-runtime] # by name: it speaks their protocols +excludes: [] +tags: [observability] # for finding it, never for resolving it + +runs: # what applying it means + - container: ... +constraints: # what must be true of a node, never which node + - architecture: amd64 +``` + +`postgres` and `lavinmq` are concrete because this module speaks their wire protocols and would +break against anything else. `container-runtime` is abstract and satisfied by the **node**. + +**`excludes`** is the one genuinely new relation: naming something that cannot coexist with +this. It is not derivable from requires and provides, and without it two modules that both +provide `terminal` — or two that both want port 443 — look independent right up until +installing the second breaks the first. + +**`constraints` are not placement.** They say what must be true of a node, never which node. +Which node runs what is an inventory decision, and the measurement found the current catalogue +deciding it in the manifest — a module pinning its database to a named node, so that a second +node cannot provide it without editing the module that consumes it. That is the mistake this +separation exists to avoid. + +## What it deletes + +The question the effort opened with, answered for the design rather than for what exists. + +- **The distinction between requiring a module and requiring a resource.** One relation. +- **The interface as a kind of thing.** A name with several providers. +- **Capability checking as a separate mechanism.** The node is a provider. +- **The domain module** — a `networking` module that exists to gather a firewall, a resolver and + a proxy under one name. It came from an older shape and does not fit: there is no such thing + to install. There is **core infrastructure**, which is a set of concrete modules named + individually — a firewall, a store, a resolver — with no flavour and no grouping module + standing in front of them. +- **Domain grouping as structure** ([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)). + Folders assert relationships; edges record them. What grouping was for — finding things, + seeing what belongs together — is a **tag** and a *query* over the graph, neither of which + anybody has to keep true by hand. +- **Tiers as a separate concept**, possibly. A tier is a level in the graph, and levels are + computed. Whether the coarse boundary is still worth naming is below. + +## What it does not delete, and should not + +**A resolver still has to exist**, and it will need version constraints, conflict handling and +a story for when two modules provide one name and nothing says which to use. That is the +long-solved and easily-botched part, and the design should say what it **delegates** to the +platform's package manager rather than reimplementing. + +## Open + +- **When two modules provide one abstract name, who chooses?** A node with both `kitty` and + `xterm` satisfies `terminal` twice. Whether the choice is a node setting, a mesh setting, or + an explicit pin is the substance of the interface idea, and this proposal does not settle it. + It is a smaller question than it was: the strict substitutability test means the consumer + genuinely does not care which it gets, so the choice is about preference rather than + correctness. +- **Are tiers a view of the graph, or a boundary that survives it?** If tiers are levels, the + concept is derived and the word is a convenience. If the tier rule — *a tier may depend only + on tiers below it* — is meant to be enforceable, it is a constraint on the graph rather than + a description of it, and it has to be stated as one. +- **What does a provider hand back?** A consumer requiring `postgres` needs credentials and an + address; one requiring `terminal` needs a command. The requirement is satisfied by the same + relation in both cases, and what flows across it is not the same shape. Whether that lives in + the name, beside it, or in what the provider returns is undecided. +- **What does a node provide that is not a capability?** Its architecture is a provided name + under this scheme, and so is its operating system. That may be elegant or may be one + abstraction too far; nothing here tests it. diff --git a/01-RESEARCH/011-the-module-graph/worked-provider.md b/01-RESEARCH/011-the-module-graph/worked-provider.md new file mode 100644 index 0000000..a92fc08 --- /dev/null +++ b/01-RESEARCH/011-the-module-graph/worked-provider.md @@ -0,0 +1,491 @@ +# A provider, all the way through + +One module, worked out completely, because it is the case that breaks the tidy version. It +looks like *a container that runs a database* and it is at least nine things. + +Postgres is the example. **The shape is not specific to it** — see the end. + +## What it carries + +**1 — A supervised container.** An image, a version, a data volume, and configuration. Easy, +and the only part the phrase "a docker service" describes. + +**2 — Persistent state, and where it lives matters.** The volume is the database. Moving this +module between nodes is not rescheduling; it is a migration. Almost nothing else in the +catalogue has this property, and nothing in `provides` / `requires` expresses it. + +**3 — Configuration that is partly the machine's.** Tuning follows the hardware — memory, +storage. A declaration generated centrally cannot know those, and +[research 012](../012-the-minimum-viable-node/00-overview.md) says the machine's own values win +on conflict. So some of this module's configuration is *derived from the node it lands on*. + +**4 — A tool surface.** It exposes capabilities agents can call — query, list, provision. That +is not a resource on a machine and not an artifact; it is a contract the mesh publishes on the +module's behalf. + +**5 — A provisioner.** The part that matters, and the one below. + +**6 — Its own bookkeeping.** The provisioner must remember what it granted to whom, or it +cannot revoke, rotate or clean up. So a module that provides state to others *also* holds state +about its providing — and that state is not the database's data. + +**7 — An exposure decision, per node it runs on.** Reachable from the machine only, from the +local network, or publicly. That is a property of *this assignment*, not of the module: the +same module on two nodes may answer differently. + +**8 — Credentials it generates.** Per consumer, and they have to reach the consumer. Which +means a provisioning edge carries a payload, and the payload is a secret. + +**9 — Health that is not "the container is up".** A container running and a database accepting +connections are different facts, and the second is the one anything cares about. This is the +host's read-back rule, at a distance. + +## The provisioner is a second kind of edge + +The tidy version of this effort said: *a module provides names, a module requires names, that is +the only edge.* A small game wanting to store data shows it is not. + +``` +my-cool-game requires postgres # I speak its protocol, it must exist +my-cool-game requires a database FROM postgres, called my-cool-game +``` + +The first is **presence**: the thing exists and is reachable. Nothing is created; nothing flows +back. `vscode requires terminal` is this, and so is `requires container-runtime`. + +The second is **instantiation**: the provider is asked to make something *for this consumer*, +and hands back what the consumer needs to use it. A database, a user, a password, an address. + +They differ in every way that matters: + +| | presence | instantiation | +|---|---|---| +| creates something | no | yes, one per consumer | +| carries a payload back | no | credentials, an address | +| can be revoked | — | yes, and must be when the consumer goes | +| provider holds state about it | no | yes — who was granted what | +| satisfied by | anything providing the name | that provider, specifically | + +**This is the mesh's actual power**, in the operator's words: a small game declares it wants a +database and the mesh makes one. Nobody creates a user by hand, nobody pastes a connection +string. That is worth being precise about rather than folding into a single relation because +one relation is prettier. + +## What that costs the design + +**The proposal's "one kind of edge" is wrong**, and the current system already knew: it has +`dependencies` for presence and `requires: provision:` for instantiation, with the resolver +deriving a presence edge from every instantiation edge. [`analysis.md`](analysis.md) recorded +that derivation as a convenience. It is not — it is the correct relationship between two +genuinely different relations. + +So: **two kinds of edge, one graph.** Instantiation implies presence. Presence does not imply +instantiation. + +## Which provider, when two nodes run one + +Asked directly, because two nodes can each run a relational store and a consumer has to be +served by one of them. Neither obvious answer is right. + +**Not "the consumer names the node."** That is placement in the consumer's manifest — a small +game edited because a database moved, which is the fault +[`proposal.md`](proposal.md) separates constraints from placement to avoid. + +**Not "the consumer does not care" either.** For presence it genuinely does not: a terminal is a +terminal. For instantiation it cares permanently, because the data lands in exactly one store +and the wrong choice is discovered long afterwards. + +**What the consumer does know is the scope of its own need.** Not which node — how many of the +thing it wants, relative to itself: + +| Scope | Means | Example | +|---|---|---| +| **shared** | one instance serves every instance of this consumer | the mesh's own registry: every node reads the same rows | +| **per instance** | each instance of this consumer gets its own | a local cache, a per-node queue | + +That is a property of the consumer, expressible without naming anything. And it is the thing +that actually decides: a shared need cannot be satisfied by a provider each node runs +separately, and a per-instance need should not be satisfied by a shared one. + +**Then the mesh binds, and the binding is written down.** Not recomputed: a resolver that +re-derives which store serves a consumer will one day derive a different answer and relocate a +database, so the binding is made once and changed only deliberately. + +**Where it is written down follows the scope.** A shared grant belongs to the *module* and every +assignment of it references the same one — which is the answer for one module installed on two +nodes wanting one database between them. A per-instance grant belongs to the *assignment*. Same +relation, two homes, and which home is not a detail: it is what makes two instances share +something or not. + +The pieces that follow, none of them settled here: + +- **When several providers satisfy the scope**, something chooses — most plausibly locality, + preferring a provider on the same node. That is a default, and it must be overridable, because + the reason to override it is exactly the reason nobody anticipated it. +- **A binding is a thing that can be wrong.** Once recorded it can be inspected, and a consumer + bound to a store on a node that no longer exists is a question somebody can be asked rather + than a failure at connect time. +- **Moving a binding moves data.** Whatever the mechanism, changing it is a migration and not a + configuration change, and a design that lets it look like the latter will lose something. + +## Provisioning is early, not late + +An assumption worth killing: that provisioning is something the control plane does for +consumers once a mesh is running. + +**The mesh's own registry is a provisioned database.** So is its virtual host on the broker. +Neither exists until something creates them, and nothing in the mesh works until they do. The +order is: + +``` +1 the store runs from the bundle the host carries +2 a database is created in it a provisioning step +3 the mesh's own schema is applied a migration, against that database +4 the control plane starts and only now is there a mesh +5 everything else is provisioned the ordinary path +``` + +Steps 2 and 3 happen **before there is a mesh to do them**. So provisioning is not a +control-plane service that consumers use; it is part of the bootstrap, and part of what the +carried bundle has to be able to express. + +**Which strains what a declaration is.** [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) +has the host applying *declared state on this machine*. A database inside a running store is not +a file or a unit — and at bootstrap it is, at least, local: the store is on the same machine as +the host applying the bundle. + +Later it is not. A consumer on one node provisioned from a store on another is the ordinary +case, and reaching it is not the host's job. + +**Resolved as two mechanisms, which is the answer rather than a compromise** +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). The host +runs bootstrap actions locally from the bundle; the control plane provisions across the mesh +afterwards. Different actors, different scopes, different trust paths — so there is no single +operation with a tier boundary running through it. + +## Several modules, one database — and the case for refusing + +Asked, then reconsidered by the operator: *maybe we should not allow it.* + +The permissive version was a per-consumer **schema** inside a shared database — its own +namespace, its own migrations, revocable by dropping the schema, with a cross-context join +possible but deliberate. + +**The stricter version is better, and it goes further than schemas.** + +> **A module is only ever granted a resource it exclusively owns.** + +No shared writes. And **no read-only role on another module's database either** — reading +another context's tables couples you to its layout exactly as firmly as writing them does, and +the coupling is harder to see because nothing breaks until the owner changes a column. + +That is what [`how-we-build`](../../00-META/how-we-build.md) §4 already says: *contexts integrate +through the record, never through a shared schema.* The permissive version kept the letter of it +and left the temptation in place, and the path of least resistance wins eventually. A boundary +that is merely inconvenient to cross is a boundary that gets crossed. + +### What it costs + +**Cross-module reporting.** Anything wanting to know what several modules hold can no longer +join across them. It consumes their events, or calls their interface, and neither is as +immediate as a query. + +That cost is the point rather than a regrettable side effect — it is §4's whole argument, and +the mesh already has both mechanisms: an event stream every context publishes to, and a tool +surface every module exposes. What gets harder is the thing that was making work belonging to +one context keep having to be implemented in another. + +**One more connection per consumer.** A dozen modules means a dozen databases rather than a +dozen schemas in one. For a relational store this is unremarkable; it is worth stating only so +nobody discovers it as a surprise. + +### What it deletes + +The effort has been looking for what the design *removes* rather than adds, and this is the +first clear instance: + +- **Grant kinds.** There is one — an exclusive resource. No schema grants, no read roles, no + scoping rules for who may see what inside a shared thing. +- **The question of who owns which table**, and with it the guessing at revocation time. +- **Cross-module migration ordering.** Two modules migrating one database need their migrations + ordered against each other. Exclusive ownership means a module's migrations are ordered only + against itself. +- **A whole class of permission modelling** that a shared store would otherwise need. + +### The rule is about contexts, not processes + +An earlier version of this file argued that a dashboard reading a dozen stores was caught by the +rule, because a dashboard is a surface and surfaces speak to an interface. **That was wrong, and +it drew the line in the wrong place.** + +The mesh's own board showing nodes, modules and deployments is not a separate context reaching +across a boundary. It is the mesh showing its own data. Reading that store is not a violation +however it is done, and requiring it to go through an interface to reach facts its own context +owns would be ceremony. + +**The rule is: a context is granted what it exclusively owns.** Everything inside that context — +its service, its surface, its tools — reads it freely. What is forbidden is a *different* +context reading it. + +Which is exactly what the count below shows: the problem was never surfaces. It was three other +contexts keeping their tables in the mesh's database. + +**And one surface over several contexts is normal.** The board visualises the mesh, the work +engine, the knowledge base and more, and the alternative — a separate web application per +context — is worse for everyone who uses it. That is not a compromise with the rule; composing +several sources into one view is what a surface *is*. + +What it changes is only where it reads from: each context's **interface**, not each context's +**store**. Most of that already exists — 56 of 126 modules carry a tool surface, more than carry +a service. + +**And the unified board is what keeps those interfaces honest.** If a view cannot be built from +a context's interface, that interface is inadequate — discovered in the one place where it is +cheap to notice, rather than the first time something else needs the same data and quietly +reaches for the store instead. + +**But "the board calls each context's interface" is not quite it either**, and the objection is +right: if every context runs its own service with its own interface, the board is coupled to N +of them instead of N schemas, something has to compose them, and composition is logic — which +tier 3 says a surface does not hold. That moves the problem up a layer rather than solving it. + +**The skeleton already answers this and the argument above talked past it.** `work` and +`knowledge` are not separate services; they are **contexts inside the control plane**, alongside +the record, inventory, delivery and the rest — and `api` is listed there as *the one interface +every surface speaks to*. + +So the board speaks to **one** interface. Behind it the contexts stay separate: separate stores, +integrating through the record. But they are one tier, one repository, one deployable, and +coupling *within* a tier is not what the tier rule forbids. + +Which resolves the objection rather than deflecting it: the problem does move up a layer, and +the layer it moves to already exists and has this as its job. + +**The caveat is load-bearing.** This holds only while the contexts are not separate deployables. +The moment one becomes its own service with its own interface, the board is back to N clients, +something must compose them, and the composition has nowhere to live that tier 3 permits. That +is a real constraint on how far the control plane may be split, and it is worth knowing now +rather than discovering it by splitting. + +If composing even one interface turns out too slow, the answer is a projection the board owns +and keeps current from events — not access to somebody else's tables. + +### Checked against the real consumers + +The rule's survival turned on whether every reader of the mesh's own registry could be served +some other way. **Eighteen consumers open a direct connection to it.** Four groups, and only one +of them is work. + +| Group | What happens under the rule | +|---|---| +| **The owner and its machinery** — the mesh module, the SDK, the environment and configuration synchronisers, secrets | Nothing. It owns the database. | +| **Node appliers** — the overlay, the shell daemon, the resolver | **Already resolved.** [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) stops the host querying the mesh database, decided for tier reasons with nothing to do with this. | +| **Foreign tenants** — the work engine (10 tables), the knowledge base (2), pipeline logs (1) | They need **their own database**. They are not reading the registry; they are storing their own data in it. | +| **Genuine cross-context reads** — the work engine reads `nodes`; two others read a handful | The only ones needing an interface or events. | + +**Thirteen foreign tables live in the mesh's registry database**, belonging to three separate +contexts. That is [`how-we-build`](../../00-META/how-we-build.md) §4's shared schema, counted. + +**So the question was the wrong shape.** The bulk of the problem is not readers needing a new +route to data — it is **tenants needing to move out**. Tasks, agents and teams have nothing to do +with nodes and modules; they are co-located by history. Give that context its own database and +its dependency on the registry shrinks to a single table. + +What remains is a handful of genuine cross-context reads, small enough to enumerate rather than +estimate. **The rule holds.** + +### Request or subscription, and what decides + +The remaining cross-context reads need one or the other. **Neither is SQL** — under exclusive +ownership a module runs SQL against its own database and nothing else, whatever transport a +query might travel over. Both options are the mesh's own channel, and both ride the broker +([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)), so the transport is +not the distinction. + +**The distinction is where the answer lives when you need it.** + +| | request | subscription | +|---|---|---| +| you ask | at the moment you need to know | never — you are told | +| the answer lives | on the other side | in your own store | +| freshness | always current | as current as the last event you received | +| when the other side is down | you cannot answer | you answer from your copy | +| what you must handle | a round trip that can fail | events you missed while you were down | + +**What decides is not taste.** [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) +makes disconnection an ordinary situation rather than an exception. So: + +> **Anything that must keep working while disconnected cannot use a request** — there is nobody +> to ask. It needs a local copy, which means a subscription. + +And the converse: anything that must be *correct at the instant of asking*, where a stale answer +is worse than no answer, cannot use a subscription. A display can lag. A decision about whether +a grant is still valid cannot. + +That turns an apparently open question into a derived one. What remains genuinely open is +narrower: **what a consumer does about the events it missed** while it was disconnected — replay +from a point, ask once for a full picture and resume, or rebuild from scratch. That is the same +question every projection has, and nothing in the record answers it yet. + +## A migration belongs to the consumer and runs on the provider + +A game defines migrations. They run against the database the store granted **it**. So the +migration is: + +- **owned** by the consumer — it is that module's schema, versioned with that module; +- **hosted** by the provider — it runs inside something the consumer does not control; +- **ordered** after the provisioning edge — there is nothing to migrate until the grant exists; +- **scoped** to the grant — the consumer's migrations touch its database and no other. + +Ownership crosses the edge, which nothing in *provides* and *requires* expresses. And it is the +*action* category from [`features.md`](features.md) made concrete: not an artifact, not node +state, and not something the host can apply, because the thing it changes is not the machine. + +It also gives a consumer's own install an internal order — **provisioned, then migrated, then +started** — that depends on an edge rather than on the module's contents. + +## What still has no answer + +**How many instances of postgres should exist?** One per mesh is wrong — a node that must work +while disconnected cannot depend on a database elsewhere. One per node is wrong — the mesh's own +registry is one thing, not one per node. So the answer is per-module, and nothing in the schema +says it. This is [`cases.md`](cases.md) axis *how many instances*, and postgres is the case that +proves it cannot be a global rule. + +**What happens to a grant when the consumer is removed?** The game is uninstalled. Its database +still exists, holding its data. Dropping it silently is data loss; keeping it forever is a leak. +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) says +the host removes what it applied and no longer declares — but this is not on the host, it is +inside another module's state, and the same reasoning does not obviously carry. + +**Where does node-derived configuration come from?** (3) The control plane composes a +declaration, and cannot know this machine's memory. Either the host fills in a blank the +declaration leaves — which makes the host decide something, against +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) — or the control +plane reads the node's inventory first and composes with it. The second is consistent and means +a declaration is composed *per node from what the node reported*, which is a stronger claim than +anything recorded so far. + + +## Providing is not a substrate thing + +The four pinned services are the obvious providers, and they are not a category. + +| Service | What a consumer asks it for | +|---|---| +| a relational store | a database, a user, credentials | +| another relational store, different vendor | a database — **and not the same one** | +| a message broker | a virtual host, a user, permissions | +| an object store | a bucket and keys | +| an image registry | a repository | +| an identity provider | a client, a realm, a secret | +| an analytics service | a site, and a tracking identity | +| a low-code data platform | a base, and a token | +| an application platform | a project, which is several of the above at once | +| a mail server | a mailbox, an alias, credentials | + +**Any hosted service can be a factory.** Providing is a facet a module may have, not a kind of +module it is — which is the same conclusion the effort reached about services and applications, +arriving from the other direction. + +That kills the last reason to keep *provider* as a category. A module runs something, or grants +something, or both, or neither. + +### Two stores, and why `database` still is not a name + +Two relational stores from different vendors both grant *a database*. They are the sharpest +possible test of the substitutability rule from [`proposal.md`](proposal.md), and they fail it +completely: different wire protocol, different dialect, different driver, different client +library compiled into the consumer. + +A consumer declaring `requires: database` and being handed either would break against one of +them. So the name promises what no provider delivers — and now with two real providers in the +catalogue rather than a thought experiment. + +*Database* remains a **tag**. It is how a person finds both. It is not how a consumer names what +it needs. + +## The same shape, three more times + +The message broker has all nine. So does the object store, and so does the image registry. They +differ in what a consumer asks for — a database, a virtual host, a bucket, a repository — and in +nothing structural. + +**A substrate service is a service plus a factory.** That is the whole pattern, and there are +four of them. It generalises past the substrate too: anything that grants something per consumer +has this shape, and anything that does not is the simpler case. + +But two things differ *between* them, and both matter more than the similarity. + +### The broker cannot be managed over the broker + +[ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md) makes the broker the +channel every node takes work from, and +[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) makes it the security +boundary — everything a node applies arrives through it. + +So the module providing the broker is also **the way modules are managed**. A declaration cannot +be delivered to it over itself, and reconfiguring it is done through the thing being +reconfigured. Nothing else in the catalogue has that property; the store is consumed by the +control plane but is not how the control plane *reaches* anything. + +This is exactly what the carried bundle exists for +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)): the broker is raised from +what the host carries, before there is a channel, because there is no other way to raise it. +Recorded here because it is a constraint on *one module*, not a general rule, and a schema with +no way to say so hides it. + +### Two modules of identical shape want different instance counts + +The broker is one per mesh — a single point of failure and a single point of trust, by decision +rather than by accident. The store cannot be: a node that must keep working while disconnected +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) cannot depend on a database +somewhere else. + +Same nine properties, opposite answers. Which settles something the cases file left open: **how +many instances is not derivable from what a module is.** It is a decision per module, it has to +be declared, and nothing in `provides`, `requires` or `excludes` says it. + +### And revocation differs in consequence + +Revoking a database leaves data behind until something drops it — a leak, and recoverable. +Revoking a virtual host drops whatever had not been delivered — not recoverable, and silent. + +The relation is the same and the blast radius is not, which is an argument for the provider +deciding what revocation means rather than the mesh applying one rule to all of them. + + +## The assignment is a third thing + +Recorded because the operator tried the alternative and abandoned it: **modules were once +node-agnostic**, and it did not survive contact. + +The worked example says why. Several of a provider's nine properties are not properties of the +module at all: + +- **where its state lives** — a volume on a particular machine; +- **how it is reached** — the same module on two nodes may answer locally, on the network, or + publicly, and that is a per-assignment decision; +- **configuration derived from the hardware** — tuning follows the memory and storage of the + machine it landed on; +- **whether this instance is the one** a given consumer is provisioned from. + +None of those belong in the catalogue, because they differ per node. None belong in the node, +because they are about this module. **They belong to the pairing**, and a design with only +modules and nodes has nowhere to put them — which is what "node-agnostic" ran out of. + +So there are three entities, not two: + +> **a module** · **a node** · **an assignment**, which is a module on a node and carries its own +> configuration + +The current system already has this, arrived at the same way: environment values are stored per +module *and per node*, so a module's settings differ between the machines running it. + +**This does not put placement back in the manifest.** A module still says what must be true of a +node and never which node ([`proposal.md`](proposal.md)). What changes is that the *result* of +placing it is a thing with its own state, rather than a fact recorded on one of the two ends. + +And it makes the composed declaration question from above answerable: a declaration is built +from the module, the node's inventory, and the assignment between them. Three inputs, which is +why two were never enough. diff --git a/01-RESEARCH/012-the-minimum-viable-node/00-overview.md b/01-RESEARCH/012-the-minimum-viable-node/00-overview.md new file mode 100644 index 0000000..50fccf4 --- /dev/null +++ b/01-RESEARCH/012-the-minimum-viable-node/00-overview.md @@ -0,0 +1,199 @@ +--- +status: active +initiated: 2026-08-26 +touches: + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0011-managed-files-are-generated-never-edited.md + - 03-DESIGN/01-to-be/05-the-node-host.md + - 03-DESIGN/00-as-is/05-runtime-and-installation.md + - 01-RESEARCH/011-the-module-graph/00-overview.md +--- + +# 012 — The minimum viable node, and adopting what is already there + +## What is being investigated + +Two questions that turn out to be one: + +**What is the bare minimum to run a one-node mesh?** Not the tiers as asserted, but the actual +closure — take the thing that must run, walk what it needs, and the set that comes back is the +answer. + +**And how does a machine that is already in use become that?** A candidate node is not empty. It +has a package manager, probably a container runtime, possibly a git installation, each with +configuration somebody chose. The mesh must **own** those, and owning is not the same as finding +them present. + +## Why + +Building tier 0 reached a wall that looked like a packaging problem and is not. + +The host can be told to run a container or install a package. Both need a file — an image, an +archive — and the question was where the host gets it. That framing produced a bad trilemma: +carry everything in the bundle, download at apply time, or have something push the files in +first. Downloading fails on the first node, which cannot fetch the image registry from the image +registry it is trying to start. + +> **Qualified by [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md).** The +> reframing below still holds for what a *tailored installer* contains — the missing pieces for a +> given machine. It does **not** have to hold for container images: the installer fetches those +> by digest, because a real machine has a network and the sealed case is the lab. + +**The reframing:** the machine is not offline. What matters is *when* the fetching happens. Move +it from apply time to **build time** — build the installer on a machine that has a network, +tailored to the target, and apply it on a target that then needs nothing. That is the same move +the lab already made for its router image, and the same property the delivery design already +claims: what ships is self-contained and a deploy touches no network. + +Which makes the interesting question not *where do artifacts come from* but **what is missing +from this particular machine**, and that needs both of the questions above answered. + +## What tailoring implies + +- **The binary stays generic; the payload is tailored.** One static host per architecture. What + is machine-specific is the bundle it applies. Rebuilding the host per machine would buy + nothing. +- **Detection is the input, not a report.** What the host already reports about a machine — + its profile and inventory — is what the difference is computed against. This is the first use + of stage 1 by something other than a person reading it. + +## Adoption + +**Simply having a package installed is not enough.** If the mesh manages the container runtime, +it decides that runtime's configuration; a runtime found already installed carries settings +somebody chose, and those cannot be discovered by noticing the binary exists. + +So a machine already in use is **adopted**: what is there is read, taken over, and thereafter +generated. + +**This was the original path.** [`00-as-is/05`](../../03-DESIGN/00-as-is/05-runtime-and-installation.md) +records adoption of a pre-existing machine's configuration as the original mechanism, since made +legacy and explicitly out of scope for the lab. It returns here for a different reason than it +was dropped for, which is a thing to notice rather than to gloss. + +**It creates a state that does not exist today.** [ADR 0011](../../02-DECISIONS/0011-managed-files-are-generated-never-edited.md) +has managed files generated and never edited; adoption needs a one-time import before that rule +starts applying. Three states, and the middle one is new: + +> unmanaged → **adopted once** → generated + +**And it crosses a boundary just drawn.** [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) +says the host never touches what it did not create — the rule that stops a converger deleting +what the mesh never put there. Adoption is the deliberate act of taking ownership of exactly +that. The rule needs a companion rather than an exception: *never, unless adoption made it the +host's*, with adoption being explicit, recorded, and visible in what the host says it owns. + +## Nothing is taken over without keeping what was there + +**Before adoption touches a file, the original is kept.** Adoption happens on machines somebody +is already using, and the configuration being taken over is configuration somebody chose. A +one-way door on a working machine is not an installation, it is a risk nobody agreed to. + +This is a *never* rule rather than a courtesy, and it earns that by the same incident the mesh's +strongest rule already carries: the worst loss in this record came from a tool acting on a path +it did not own. Adoption is that act, made deliberate — which makes the safeguard obligatory +rather than optional. + +What that requires, and what remains open: where the copy lives, whether it is recorded in what +the node knows about itself so that adoption is *visibly* reversible, and whether the mesh keeps +it forever or hands it back when it stops managing the thing. + +## Adoption produces a briefing, not just a result + +Proposed by the operator, and it answers a question this effort had open with two bad answers. + +Adoption meets things a script cannot decide. A container runtime configured with one storage +driver and a mesh wanting another. A package pinned to a version somebody chose for a reason. +Local settings the mesh has no opinion about and no business discarding. Silently winning is +wrong in both directions; refusing outright makes a machine in use unadoptable. + +**So adoption has two outputs.** What it did — mechanical, recorded, in the node's state. And a +**briefing**: what it found, what it took over, and what it could not resolve, written to be +read by a person or an agent, which is the first thing a session on that node has to work with. + +Conflicts are **flagged, not resolved**. That is the same principle the declaration parser +already applies — name every problem at once, to somebody who can act on it — at a larger +scale, and applied to a case where refusing wholesale would be worse than proceeding. + +### On conflict, the machine's configuration is kept + +Decided by the operator, after first deciding the opposite — recorded that way because the +reasoning for each direction is the useful part. + +Where the existing configuration and the mesh's disagree, **what is already on the machine +stays**, the conflict is flagged, and it is reconciled afterwards. Adoption always **completes** +— flags inform, they do not block — and *adopted with open questions* prevents nothing. The +node is a node. + +**What this buys.** Adoption becomes non-destructive by construction. The class of conflict that +made the opposite rule dangerous — a storage driver against the filesystem it is actually on, a +data directory pointing at a mount that exists — cannot arise, because nothing tied to the +machine's physical reality is ever overwritten. A machine in use keeps working exactly as it +did. + +**What it exposes, which is the mirror of what it fixes.** The mesh's configuration is not only +preference. Some of it is what a module needs in order to function at all. Keeping the machine's +version there produces a module that is installed and does not work — *an installed package is +not a capability* +([04-ISSUES/007](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)) +arriving from a direction that issue did not anticipate. And a fleet where every node kept its +own settings is a fleet where a module works on one node and fails on another with nothing in +the mesh able to say why. + +**The distinction that dissolves both rules.** Neither direction is right as a blanket, because +the question is not *whose configuration wins*. It is whether the module **requires** the +setting or merely **prefers** it — required contradictions cannot be kept without breaking the +module, and preferences should always yield to what is already there. + +That is a property of the module's own declaration rather than of the adoption algorithm, which +makes it one more thing the graph would carry +([research 011](../011-the-module-graph/00-overview.md)). Until modules can say which of their +settings are load-bearing, adoption is choosing a default in the dark, and the default chosen +here is the one that does not break the machine it is adopting. + +### The briefing carries an outcome, and the outcome is derived + +Proposed by the operator: the report states plainly whether adoption succeeded, partly +succeeded or failed, and each line carries its own severity. + +**Each line is marked, and the overall is the worst mark present.** Derived rather than stated +alongside, because two fields written independently drift — and a briefing reading *full +success* while carrying a failed line is exactly the fault this record keeps cataloguing. An +outcome computed from its lines cannot disagree with them. + +| Mark | Means | +|---|---| +| `ok` | done, and verified | +| `kept` | a disagreement; the machine's value was kept and somebody should look | +| `unknown` | could not be determined | +| `failed` | could not be done — the node is not what was asked for | + +**`unknown` is not a shade of success.** Adoption will meet configuration it cannot parse and +state it cannot read, and folding those into *fine* is the same move as reporting an installed +package as a capability. A thing nobody could determine is a thing nobody can rely on, and it +gets its own mark for the same reason a capability detector reports *why*. + +**And this opens something the earlier rule did not cover.** *Flags inform, they do not block* +was decided about **conflicts** — where the mesh chose, deliberately, and the machine still +works. A **failure** is different in kind: not *we chose* but *we could not*. Treating both the +same makes a node where something the mesh needed never happened indistinguishable from one +where a log level differed. Whether a failed line still lets adoption complete is therefore +reopened by adding severity, and is not decided here. + +## Open questions + +| Question | Why it is open | +|---|---| +| What is the closure for a one-node mesh? | The skeleton asserts four pinned services. [Research 006](../006-mesh-from-scratch/00-overview.md) already asks whether it is four or five and does not answer. A graph gives a computed answer instead of an asserted one, which is [research 011](../011-the-module-graph/00-overview.md). | +| Is "tier" the same thing as a graph level? | Tiers were named as a bootstrap order. If the closure is computed, tiers may be a derived view of the graph rather than a separate concept — or they may be a coarser boundary that survives for a different reason. | +| ~~What happens when existing configuration contradicts what the mesh needs?~~ | **Answered** — the machine's configuration is kept, the conflict is flagged, and it is reconciled afterwards. | +| ~~Do flags block, or only inform?~~ | **Answered** — they inform. Adoption always completes, and the node is a node. | +| Can a module say which of its settings are load-bearing? | The question that dissolves the conflict rule rather than choosing a side. A setting the module *requires* cannot be kept from the machine without producing something installed and broken; a setting it merely *prefers* should always yield. Until a module can say which is which, adoption is defaulting in the dark. Belongs with the graph. | +| Does a `failed` line still let adoption complete? | *Flags inform, they do not block* was decided about conflicts, where the mesh chose and the machine works. A failure is *we could not*, which is different in kind — and treating them alike hides the worse one behind the commoner one. | +| How is a flagged conflict reconciled, and by whom? | The briefing hands it to a session. What that session is empowered to change, and whether the resolution is recorded so the next adoption does not re-raise it, is undecided. | +| Where does the kept original live, and for how long? | Whether it is recorded in the node's state so adoption is visibly reversible, and whether it is returned when the mesh stops managing the thing. | +| What shape is a briefing? | Structured enough to be acted on, prose enough to be read. It is the first thing a session on a new node sees, which makes it an interface rather than a log. | +| Does owning a package mean owning its version? | Owning configuration and owning the package are different scopes. The second means the mesh decides which version is installed, and that decision then has to survive the machine's own package manager updating it. | +| How does a bundle stay true between building and applying? | It is built against a scan of the target. The machine can move between the scan and the apply, so the host has to verify rather than assume — and fail plainly when the bundle no longer fits. | +| What cannot be precomputed at all? | Anything built from source on the target still needs a toolchain and a network at that moment. Tailoring moves that cost rather than removing it, and *minimal viable* has to be honest about what it cannot ship ahead. | +| Does presence differencing understate the gap? | Knowing a package manager is installed does not say it is the version the mesh needs. A difference computed on presence alone is optimistic. | diff --git a/02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md b/02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md similarity index 77% rename from 02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md rename to 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md index 3ad2729..546301e 100644 --- a/02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md +++ b/02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md @@ -1,11 +1,12 @@ --- +topic: the mesh status: accepted date: 2026-08-22 deciders: jochen reconstructed: false --- -# 15. The mesh brokers capabilities; nodes host; agents think +# 1. The mesh brokers capabilities; nodes host; agents think ## Context @@ -69,6 +70,52 @@ invariants were found violated simultaneously (see Consequences). **The mesh brokers capabilities. Nodes are places where work runs. Agents are personas that think and act.** Everything else supports one of those three. +### What this is, plainly — and what "mesh" does not mean + +*Written 2026-08-29, from working through connectivity and asking whether the word still fits.* + +Four layers. Naming them honestly is worth more than the word on the tin: + +| | | +|---|---| +| **machines are linked by a private network** | and every machine reaches every other over it | +| **one node holds knowledge of all of them** | the control plane, and only it | +| **modules are how anything is built and delivered** | this *is* the CI/CD, not something beside it ([ADR 0010](0010-delivery.md)) | +| **every node runs a session you can message** | a feature of the node, remembering across callers; any node can message any other ([ADR 0004](0004-a-node-and-how-it-joins.md)) | +| **workers are hired onto nodes to do tasks** | employees, with a lifecycle — a different thing from the row above ([ADR 0003](0003-agents-are-persistent-employees.md)) | + +**The last two rows are not the same thing and the vocabulary of one does not describe the other.** +A node's session comes with the machine: nobody hires it, it holds no tasks, it is never +reassigned, and it goes when the node leaves. A worker is an employee — named, hired, drained, +retired, movable. They are built from the same parts and run on entirely different terms, and +collapsing them is how the employee vocabulary ends up stretched over something it does not fit. + +**Neither makes the node itself a thinking thing.** A node is a machine; both of these run *on* +one, which is why *a node does not authenticate to a model provider, agents do* is unaffected by +either. + +**Where the value is** is that both can reach across the whole set: a shell, a service, a file, or +simply a question to another node. Not machines that can be configured centrally, which is +ordinary, but a set of machines that can be worked across as though they were one. + +**This is not a mesh in the peer-to-peer sense and will not become one.** The word describes what +machines can reach, not how they are governed: + +| | a mesh? | +|---|---| +| what a machine can reach | **yes** — genuinely any to any | +| how the traffic travels | no — anything crossing sites transits the hub | +| who decides | no. One node, declared | + +**And *master* overstates it in the other direction.** A master implies the others need it in +order to function. They do not: every node holds what it was last told and runs from that copy +*always* — not as a fallback, as the only mode it has. So the control plane being gone is every +node in the ordinary disconnected situation at once, and **what is lost is change, not +operation.** + +The accurate phrase is **one authority, no failover**, and both halves are deliberate +([ADR 0006](0006-the-substrate-and-the-control-plane.md)). + ### Nodes and agents are decoupled A node is a place where an agent can run — that is the entire relationship. There is no @@ -182,7 +229,7 @@ existing pipeline. Nothing here requires a flag day, and nothing here is cheap. - `modules/hal/sdk/src/feature-handlers/index.ts` — `FEATURE_HANDLERS`, the fixed handler array that makes a feature a singleton per module - `modules/postgres/tools/index.ts` — the adoption path that rotates a shared credential -- Mediahuis `papa-hq`, ADR 0009 *Composable, independently-shippable modules* — the +- Mediahuis `papa-hq`, ADR 0020 *Composable, independently-shippable modules* — the constraints that make a unit independently shippable, applicable unchanged to features - impire.io / soulstream — *the record* as integration substrate, personas over services, and "cheap awareness and expensive thinking" diff --git a/02-DECISIONS/0002-everything-is-a-module.md b/02-DECISIONS/0002-everything-is-a-module.md deleted file mode 100644 index 0811935..0000000 --- a/02-DECISIONS/0002-everything-is-a-module.md +++ /dev/null @@ -1,73 +0,0 @@ ---- -status: accepted -date: 2026-03-14 -deciders: jochen -reconstructed: true ---- - -# 2. Everything is a module, and one manifest describes all of them - -> Reconstructed after the fact from the evidence cited below. - -## Context - -The mesh carries several kinds of thing: containerised services with data and ports, pure -capability providers with no service at all, and bare markers whose only content is that a -node has them. Before this decision these were separate concepts with separate handling — -the earlier vocabulary was *capabilities*, and services were installed by a different path -than tools. - -Every distinct kind of thing needs its own install path, its own change detection, its own -place in the delivery pipeline, and its own documentation. Three kinds means three of each, -and every new feature has to be built three times or, more commonly, once — leaving two kinds -quietly unsupported. - -## Considered options - -1. **Separate concepts per kind** — a service registry, a tool registry, a node feature flag - list. Rejected: it is what existed, and the cost was paid in every cross-cutting change. -2. **One manifest, kind inferred from directory contents.** Chosen. -3. **One manifest with an explicit `type:` field on every module.** Partly adopted — a service - still declares itself — but the general rule became inference, because a declared list and - the directory it describes drift, and the directory is the one that is true. - -## Decision - -Everything the mesh installs is a **module**: a directory with a manifest. The manifest -declares identity, environment variables, what the module provides, what it requires, and how -it is exposed. What kind of module it is follows from what the directory contains: - -| Contains | Is | -|---|---| -| a compose definition | a service | -| a tools directory | a capability provider | -| a daemon or unit directory | a long-running process | -| a configs directory | a source of managed files | -| nothing but a manifest | a flag — presence is the whole content | - -A module may be several of these at once. Each is a **feature**, and the delivery pipeline -addresses features, not modules. - -The mesh's own components are modules on exactly these terms. They get no privileged install -path, no separate registry, and no exemption from the pipeline. - -## Consequences - -- One mechanism to learn, one to document, one to fix. A pipeline improvement reaches - everything the mesh carries. -- Dogfooding stops being a discipline and becomes structural: if the mesh's own components - need an exception, the machinery is unfinished, and that is visible immediately. -- Feature detection from directory contents means a directory rename silently changes what a - module *is*. This has bitten repeatedly — a hook named for a feature the module does not - have is skipped without complaint. -- The manifest becomes load-bearing and grows. It is now the largest single point of - coupling in the mesh. - -## References - -- `Rename capabilities → modules across the entire codebase`, 2026-03-14. -- `Merge fail2ban, ufw, firewall apps into modules`, 2026-03-15 — the first modules to arrive - by conversion rather than by creation. -- Knowledge base: `modules`, `modules/manifest-reference`, `conventions/modules`. -- The rename-breaks-detection shape: `troubleshooting/hooks-named-for-missing-feature`, - `troubleshooting/health-check-tools-index-false-positive`. diff --git a/02-DECISIONS/0001-nodes-communicate-over-a-broker.md b/02-DECISIONS/0002-nodes-communicate-over-a-broker.md similarity index 97% rename from 02-DECISIONS/0001-nodes-communicate-over-a-broker.md rename to 02-DECISIONS/0002-nodes-communicate-over-a-broker.md index 3a841e3..b548c9b 100644 --- a/02-DECISIONS/0001-nodes-communicate-over-a-broker.md +++ b/02-DECISIONS/0002-nodes-communicate-over-a-broker.md @@ -1,11 +1,12 @@ --- +topic: the mesh status: accepted date: 2026-02-25 deciders: jochen reconstructed: true --- -# 1. Nodes communicate over a message broker, not over HTTP +# 2. Nodes communicate over a message broker, not over HTTP > Reconstructed after the fact from the evidence cited below. The decision was taken in > implementation, not in a record; this document states what was decided and why, not a diff --git a/02-DECISIONS/0012-agents-are-persistent-employees.md b/02-DECISIONS/0003-agents-are-persistent-employees.md similarity index 96% rename from 02-DECISIONS/0012-agents-are-persistent-employees.md rename to 02-DECISIONS/0003-agents-are-persistent-employees.md index 2530e82..5bbe516 100644 --- a/02-DECISIONS/0012-agents-are-persistent-employees.md +++ b/02-DECISIONS/0003-agents-are-persistent-employees.md @@ -1,11 +1,12 @@ --- +topic: the mesh status: accepted date: 2026-07-12 deciders: jochen reconstructed: true --- -# 12. An agent is a persistent employee, not an instance of a pool +# 3. An agent is a persistent employee, not an instance of a pool > Reconstructed after the fact from the evidence cited below. @@ -59,7 +60,7 @@ itself is an agent of a kind exempt from the hiring lifecycle. or another agent is hired — both deliberate acts. - The transition was not free. Lifecycle columns had to reach every query that selects an agent, and the ones that were missed failed at the moment of hiring rather than at startup. -- This is the decision [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) generalises: +- This is the decision [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) generalises: one kind of participant, differing only in modality. ## References diff --git a/02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md b/02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md deleted file mode 100644 index 1b91bcc..0000000 --- a/02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md +++ /dev/null @@ -1,65 +0,0 @@ ---- -status: accepted -date: 2026-04-02 -deciders: jochen -reconstructed: true ---- - -# 3. The mesh database is the source of truth; the repository is node-agnostic - -> Reconstructed after the fact from the evidence cited below. - -## Context - -Two things must be known to run the mesh: **what exists** — which modules there are, what each -declares, how each is built — and **what runs where** — which node hosts which module, with -which settings, at which version. - -The repository is the natural home of the first. It was initially also the home of the second: -per-node directories held that node's configuration, and adopting a machine meant committing -its files. That has three costs. A node cannot be changed without a commit, so runtime state -and source share a review cadence they do not share a rhythm with. Two nodes cannot be -reconciled, because nothing holds both. And the repository becomes an inventory of the -installation, which is exactly the content that cannot be made public. - -## Considered options - -1. **Per-node directories in the repository.** Rejected — it is what existed. Every binding - change is a commit and a deploy, and the repository accumulates an inventory of one - particular mesh. -2. **Configuration files distributed to nodes and edited there.** Rejected. There is then no - authority: two nodes disagreeing have no arbiter, and drift is invisible until something - breaks. -3. **A mesh database as the single authority, cached locally for resilience.** Chosen. - -## Decision - -A single database holds every binding: which node hosts which module, at which selection, with -which environment overrides, plus mesh-level settings that all nodes read. The runtime loads -its configuration from that database at startup and falls back to a local cache when the -database is unreachable. - -**The repository defines what exists. The database defines what runs where.** No node-to-module -mapping is ever committed. - -A node is therefore not described anywhere in source. Bringing one into the mesh is a database -operation. - -## Consequences - -- The repository becomes node-agnostic, and can be published without disclosing an - installation. This repository's public stance rests on that property. -- A binding changes without a commit, a build, or a deploy. -- The local cache means a node survives losing the database, but a node running from cache is - running from a snapshot with no indication of its age. Divergence is silent by construction. -- The database is the hardest dependency in the mesh. It is also a module, provisioned like - any other, which makes its bootstrap circular — resolved by the first-node initialisation - script, and the reason such a script exists. -- Nothing on a node is authoritative. That is what makes the next decision necessary. - -## References - -- `Phase 3: rename core modules to hal/ namespace`, 2026-04-02, and the mesh configuration - tables that landed with it. -- Knowledge base: `mesh` — "The repo is node-agnostic. It contains no per-node assignments." -- The stale-cache shape: `troubleshooting/installed-version-and-deployments-are-stale`. diff --git a/02-DECISIONS/0004-a-node-and-how-it-joins.md b/02-DECISIONS/0004-a-node-and-how-it-joins.md new file mode 100644 index 0000000..a958a76 --- /dev/null +++ b/02-DECISIONS/0004-a-node-and-how-it-joins.md @@ -0,0 +1,263 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-28 +deciders: jochen +reconstructed: false +--- + +# 4. A node, and how it joins + +*Consolidated 2026-08-28 from four records.* + +## What a node is + +**A managed machine inside the mesh.** Not a device that is merely known about, not an +unprivileged something. If the mesh does not manage it, it is not a node — it is a client, a +peer, or a thing on the network, and those want their own names rather than a weakened version of +this one. + +**A disconnected node is still a node, in a different situation.** Reachability is **state, not +class**. A machine switched off, roaming, or behind a connection that has dropped has not become +a lesser kind of thing; it has a last-known state and a pending set of declarations. + +The distinction people reach for is real, but it is **capability** — what this machine can be +asked to do — and that belongs in the host's profile rather than in the definition of a node. + +**This is the rule that does the most work elsewhere.** A single control plane is tolerable +because its absence is every node in the ordinary disconnected situation at once. An episodic +host on a phone is that situation more often. Neither needed a new mechanism. + +### A node runs one agent session + +*Written 2026-08-29. It runs on every node today and appeared in no record, which is how something +deliberate comes to look accidental.* + +**A node is a machine. The session is a feature of it** — one of the things running there, like the +host, like any workload. The node does not think; something on the node does. Which is why +[ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)'s *a node does not authenticate to a model +provider, agents do* holds unchanged: the session authenticates, and it is not the machine. + +**It is permanent, and it remembers.** Anything in the mesh can send it a message; it replies; and +what it was asked ten minutes ago is still there next week, alongside what everything else asked in +between — the same way both sides of any conversation remember it. + +Its system prompt is the node's **engram** — what makes one node's replies recognisably its own +rather than generic. + +**It has its own tools**, and fewer than a session a person is driving directly. So a question can +be answered by going and looking: *what is in our forge*, not only *what is your battery*. + +**Messages travel the broker like everything else** ([ADR 0002](0002-nodes-communicate-over-a-broker.md)). +There is no second transport and nothing is dialled. + +**Any node can message any node, and this is the one part of the system that is genuinely a mesh** +— symmetric, with no centre. A node that is asked something it does not have can ask another, and +how it passes the question on is its own business: it may say who wants to know, or simply ask. A +person relaying a question makes the same choice, and it follows from the engram rather than from a +message format. + +**There is no authorisation between nodes.** Every node is the operator's own, so a message from +one is a message from them, and asking a node something is asking a colleague rather than +presenting credentials. Stated once so it is not discovered later: **the mesh boundary is therefore +the security boundary** — anything inside can reach what any node can reach, which is what puts the +whole perimeter on the token and the overlay +([ADR 0007](0007-connectivity.md)). + +**It can be switched off, and switched off it still answers.** A node whose session is disabled +replies saying so, at once, with no model involved — the queue is still read, and the state is the +reply. That is deliberate and it is the same rule the host follows about a service that does not +exist: **absence must never be indistinguishable from a failure to answer.** A node with nothing +there is a silence somebody has to go and diagnose; a node that says *I am switched off* is not. + +**One per node, always, and it cannot be moved to another machine.** Two and nothing decides which +replies; none and the node is mute; moved, and one machine is answering as another. + +**It is not an employee** ([ADR 0003](0003-agents-are-persistent-employees.md)). Nobody hires it, +it holds no tasks, it drains nothing and it is never reassigned — that vocabulary was written for +workers and does not describe this. It exists because the node does, and it is gone when the node +leaves. + +## How it joins + +**The host has one behaviour and two sources of declaration.** What differs between the first +node and the fiftieth is not what the host does but where the declaration comes from — and, as +above, that is a situation rather than a class. + +| | declaration comes from | +|---|---| +| no mesh reachable | the pinned bundle the host carries | +| mesh reachable | the control plane, over the link | + +**The first node is not a different kind of node.** It is a node whose mesh is not up yet. It +applies the bundle it carries, the control plane comes up on top of it, and from that moment it +takes declarations like everything else. **Its specialness is temporary and self-erasing**, which +is what the hand-run bootstrap scripts never were. + +**A joining node does the minimum to be reachable and nothing else** — an identity, an address, +and one peer to reach. It does **not** compute the overlay: the whole peer set is derived +centrally and pushed down. + +That is also why the migration is smaller than it looked. The hard part of the overlay — every +node's key, address, site and reachability — is only needed to compute the *whole* mesh, and a +joining node needs one peer. + +## The link is the security boundary + +**Everything reaching a node arrives one way**, and four properties make that a boundary rather +than a pipe. + +**It is outbound and node-initiated.** The node dials the control plane; nothing dials a node. Not +only defensive — most nodes sit behind a household connection with no forwarded port, so an +inbound control channel would work for one node and not the rest, and the difference would be +invisible until it mattered. **A node has no listening control surface at all.** + +**A node holds its own identity and nothing else.** No shared secret, no credential to anything it +does not own. **Compromise of a node is compromise of that node** — which the current arrangement +does not have, because every node permanently holds the same database and object-store +credentials, and there is no mechanism that rotates one and informs everything holding it. + +### What that identity is: a keypair the node generates + +*Written 2026-08-29. This is the same rule as the sentence above, and it had been treated as an +open question for weeks because of a word.* + +**The node generates a keypair. The private half never leaves the machine. The mesh records the +public half.** Ed25519, the same as the control plane's signing key, in the other direction: +the mesh proves itself to a node by signing, and a node proves itself to the mesh by signing. + +**This was never open.** [`08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md) already +says it of the overlay keys, in these words: *each node generates its own keypair, the private key +never leaves the machine, the public key is published to the mesh* — and adds that this **is** +ADR 0004's *a node holds its own identity*, applied. What was missing was applying it to the thing +this record is about. + +**The word that caused it:** the lifecycle says a joining node *receives* its own durable identity, +which reads as the mesh issuing something, and then the question becomes *issuing what*. It does +not issue anything. The node arrives holding its identity; what it receives is **being known**. +Enrolment is the moment the mesh writes down a public key it will believe, and the one-time secret +is what buys the right to have it written down. + +**Everything above then holds literally.** Nothing is stored that could be stolen and replayed: the +mesh's copy is a public key, so a copy of the mesh's database grants nothing. *Compromise of a node +is compromise of that node* becomes true rather than aspirational, because the only secret on a +machine is the one that identifies it. + +### What "connecting to the mesh" is, concretely + +*Written 2026-08-29, because it was asked and this record had never said it.* + +**One outbound AMQP connection from the node to the broker, held open.** That is all of it. There +is no second connection and nothing is ever dialled *at* a node. Being in the mesh, operationally, +means that connection is up; being disconnected means it is not +([`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md)). + +**Two different things ride on it, and conflating them is what made this confusing:** + +| | what it answers | who issues it | +|---|---|---| +| **an AMQP account** | may this connection be accepted at all | **the mesh, at enrolment** | +| **the node's keypair** | which node is speaking, on every message | **the node**, above | + +**The account is the mesh's to issue**, and per node. The broker has to authenticate somebody +before a connection exists, and a shared account would let any node consume another's queue — +which is the shared-credential fault this record exists to remove, reappearing at the transport. +So enrolment creates that node's account and hands it over, and it is rotatable without touching +the node's identity. + +**The keypair is not made redundant by it.** With only an account, the control plane would know +which node is speaking *because the broker says so* — and that is the same transitive authority +this record refuses in the other direction. A compromised broker could then attribute reports to +whichever node it liked, and the control plane would act on them. Signing is what removes the +broker from the question in both directions. + +**So a node holds two things after enrolment**: a credential the mesh issued for reaching the +broker, and a key it generated itself that the mesh only ever sees the public half of. Both are +its own, neither reaches anything else, and *compromise of a node is compromise of that node* +still holds. + +### Its own key, not the machine's SSH host key + +Reusing the host key is the obvious economy and it is refused, for reasons that are operational +rather than fastidious: + +- **It is regenerated by ordinary events.** A reinstall, an image cloned, `ssh-keygen -A` on a + rebuild — each silently un-enrols the node, and the failure appears as an authentication problem + with no cause anybody changed. +- **It is managed by something else.** Its lifecycle belongs to the machine's SSH daemon, and an + identity the mesh depends on should not rotate on a schedule the mesh does not know about. +- **Not every node has one.** A partial host has no SSH daemon + ([ADR 0005](0005-the-node-host.md)), and an identity scheme that excludes a supported kind of + node is not one. + +**The mesh should still know the host key** — it knows every node, so it can distribute host keys +the same way it distributes authorised keys +([ADR 0006](0006-the-substrate-and-the-control-plane.md)), and node-to-node SSH stops depending on +trust-on-first-use. That is the good half of the idea, kept. + +**Authority is mutual.** The node proves it may join, and the control plane proves it is the +mesh. One-way is not enough: the host applies whatever the link delivers, so a node that cannot +tell the mesh from something impersonating it will apply that something's declarations. + +**What may be pushed is bounded by form, not by trust.** Declarations of known shape, never a +command to run. Stated honestly, **this bounds form and not impact**: a compromised control plane +can declare harmful state and the host will apply it faithfully, because that is what it is for. +What the property buys is that the blast radius is *describable* — exactly what the declaration +language can express, which can be reviewed. An arbitrary command channel has no such bound. + +## The enrolment token carries the mesh + +Mutual authority needs the node to verify something before it trusts anything, and that is a +circle: verifying the mesh needs the mesh's certificate authority, and obtaining one means +trusting whoever hands it over. There is a second circle beside it — a node must reach the mesh +before the mesh has configured it, so it can resolve no mesh name. + +**Both are the same shape: a node needs a fact about the mesh before it has any trustworthy way +to obtain one.** So that fact arrives by a path other than the mesh. + +**The token carries five things**, and it is the only thing a joining node needs: + +| | | +|---|---| +| **who it is** | the name the mesh calls this machine | +| **where** | the broker's **address**, not a name — there is no resolution yet, and this is why none is needed | +| **what it is connecting to** | the fingerprint of the broker's certificate | +| **who it will believe** | the control plane's signing identity | +| **the right to join** | a one-time secret, useless once used and useless after it expires | + +*The first row was added 2026-08-30, from raising a mesh end to end for the first time.* It reads +like an oversight and is not: **the node cannot work its own name out.** The name is the mesh's, +chosen when the record was created, and the broker account the node authenticates as is named +after it — so it must be known *before* the mesh can tell the node anything. It is not a secret, +and whoever issues the token already has it. + +Without it, enrolment fails at the broker with an empty username and a message about credentials, +which points at everything except the cause. **A missing fact that surfaces as an authentication +error is worse than one that surfaces as a missing fact.** + +**Carried out of band**, by the person adopting the machine. That is what breaks both circles: +its authenticity comes from the channel it travelled, not from anything the node can check +afterwards. **Trust on first use, with the first use moved out of band** — the difference between +a pin and a guess. + +**The endpoint and the authority are two identities.** A node connects to the broker and takes +instruction from the control plane behind it. Pinning only the broker would make the control +plane's authority *transitive*, and a compromised broker could then forge declarations — which, +since the host applies whatever the link delivers, is the whole machine. So the transport is +verified once at connect, and **each declaration is verified by its signature, every time**. + +**What this settles:** the mesh's certificate authority is not a bootstrap concern — it certifies +internal names once a node is a member. Nothing needs name resolution before the link. And +nothing is placed on disk beforehand except the token, which is the first moment *a node holds +only its own identity* becomes true rather than aspirational. + +## Consequences + +- **Declarations must be signed**, and the host must tell *this is not from the mesh I joined* + apart from *this is malformed*. Rotating the signing identity is a fleet-wide operation with an + overlapping rollover, and that is the cost of not trusting the broker. +- **The token becomes security-critical**, because it carries the pin. Tampering substitutes the + mesh — which is strictly better than the alternative, where there is nothing to tamper with and + the node trusts the first answer unconditionally. +- **A rejoining node is ordinary.** There is no long-lived secret to recover, so a node that lost + its identity gets a new token. diff --git a/02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md b/02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md deleted file mode 100644 index d5eff21..0000000 --- a/02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md +++ /dev/null @@ -1,67 +0,0 @@ ---- -status: accepted -date: 2026-04-06 -deciders: jochen -reconstructed: true ---- - -# 5. Capabilities are provisioned on declaration, not configured by hand - -> Reconstructed after the fact from the evidence cited below. - -## Context - -Most modules need something another module holds — a database, a cache, a bucket, a message -vhost, an identity client. Wiring that by hand means creating the resource, creating a user, -generating a credential, putting it in the consumer's configuration, and repeating all of it -on every node the consumer runs on. - -Every step is a place to make a mistake that surfaces much later, and the credential ends up -written somewhere it can be read. - -## Considered options - -1. **Manual setup, documented.** Rejected. Documentation of a manual procedure is a - description of the mistakes people will make. -2. **A shared credential per resource type**, distributed to all consumers. Rejected: no - isolation, and rotation becomes a mesh-wide outage. -3. **Declared requirements, satisfied by the provider module.** Chosen. - -## Decision - -A module declares what it **provides** and what it **requires**. A requirement names the -provider, the resource type, optionally a name and a target node, and a mapping from the -resource's connection fields to the consumer's environment variables. - -The mesh satisfies it: a provisioner belonging to the provider creates the resource and its -credential, records the grant, and writes the mapped values as database overrides. The -synchroniser from [ADR 0004](0004-managed-files-are-generated-never-edited.md) then -materialises them. Neither the credential nor the topology is ever written by hand. - -A requirement may name a provider on another node. The grant records consumer and provider -nodes separately, so cross-node wiring is the same declaration. - -## Consequences - -- **Provisioning becomes a core concern of the mesh, not plumbing.** A module asks for a - capability; where it lives is the mesh's problem. This is the property - [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) later builds the whole domain model - around. -- Credentials are never authored, so they are never authored badly, and they are never in the - repository. -- Each consumer gets its own credential, so revocation is per-consumer. -- Rotation is where this bites. A shared secret rotated for a new consumer invalidates the - peers holding the old one, and this has taken the mesh down. The declaration model makes - granting easy and says nothing about fan-out. -- A module with no requirements skips the stage entirely, which is correct and also means the - absence of provisioning is indistinguishable from provisioning that did not run. - -## References - -- `Remove shell/ helper library; split brain into independent workspaces`, 2026-04-06 — the - provisioner daemon becomes its own component. -- `Coordinator refactor: centralize pipeline orchestration`, 2026-04-04 — the provision-then- - environment-then-start sequence becomes the coordinator's. -- Knowledge base: `provisioning`, `provisioning/requires`. -- The rotation failure: `troubleshooting/provision-rotation-invalidates-peers`, - `troubleshooting/provision-adoption-rotates-live-credential`. diff --git a/02-DECISIONS/0005-the-node-host.md b/02-DECISIONS/0005-the-node-host.md new file mode 100644 index 0000000..7afc763 --- /dev/null +++ b/02-DECISIONS/0005-the-node-host.md @@ -0,0 +1,155 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-28 +deciders: jochen +reconstructed: false +--- + +# 5. The node host + +*Consolidated 2026-08-28 from eight records. Tier 0 is one component and was decided over a +week; the reasoning is kept, the fragmentation is not.* + +## It applies; it does not decide + +**The host makes a machine match what it was told, and never works out what that should be.** + +This is the line the whole tier rests on, and it is not about privilege — it is about what a +single machine can *know*. Deciding needs knowledge the machine does not have: which nodes should +run a store, which peers belong in an overlay, whether a node has been unreachable for a week. +Anything needing a second node is the control plane's. + +The practical form: **the host never queries the mesh's database and holds no credential to it.** +Two modules in the current mesh do, and they are the reason every node permanently carries a +database credential. + +## It depends on nothing that must be installed first + +**A statically linked binary. Copy it onto a machine and run it — that is the whole +installation.** Written in Go, because the job is system-level and because a runtime that must be +installed first would make the host depend on the thing it exists to install. + +**What it needs from the machine is not a dependency in this sense.** An init is not installed; +it is what the machine already is. A package manager is the distribution. Those are what a +machine *is*, not what must be put on it before the host works. + +## It is built per operating system + +**`systemd` and `pacman` are the Arch host's implementation, not abstractions the mesh grows.** +They are not independent choices: a machine has pacman *because* it is Arch, and the package +manager, service manager and packaging format arrive together as one decision somebody made at +install time. + +``` +mesh-host-arch pacman · systemctl · a container runtime +mesh-host-alpine apk · rc-service +mesh-host-android neither — a partial host +``` + +**Abstracting them was rejected on correctness, not effort.** The service applier reads systemd's +`LoadState` to tell *not installed* apart from *stopped* — which is what stops it reporting +absence as success — and OpenRC has no equivalent. An interface spanning both must drop it, and +the lowest common denominator is exactly where that fault lives. + +**Almost all of it is shared.** The declaration vocabulary, the store, the apply loop, the +read-back discipline, the refusal model and the link are portable. Two appliers differ. + +**A host that cannot implement a shape refuses it.** Android has no package manager it may drive +and no init it may register with, so it implements `file`, `directory` and `action` and refuses +the rest — the same refusal an unknown type gets, with a different reason. Those three are the +portable floor, and they are what makes a partial host a real thing rather than a broken one. + +## It is a root service, and it never manages its own unit + +**Root**, because no useful part of the job is unprivileged: it writes under `/etc`, installs +packages, manages units and runs containers. + +**It cannot run in a container**, and the reason is decisive rather than stylistic: installing the +container runtime is a step of the bootstrap, so a host inside a container would need the thing +it exists to install. Everything above tier 0 is a container; the host is not. That split is the +tier boundary made concrete. + +**The installation owns the host; the host owns everything else.** It manages `service` resources +and its own unit is one — the temptation is obvious and it ends with a host stopping itself half +way through an apply, leaving a machine with nothing running to fix it. + +## An init is asked for one thing + +**Start this at boot.** That is all, and every init can express it — systemd, OpenRC, runit, s6. + +**Everything else is a launcher the host ships**, which supervises it: restart it when it exits, +count consecutive failures, roll back after too many, halt after that. Policy in a unit file can +only be read and hoped for; a script with a counter can be tested, and this is the one piece that +must work on a machine where the host does not. + +**The launcher does not exec the host, it supervises it** — so restarting is ours rather than the +init's. The cost is signals: a supervisor that exits while its child runs leaves the host to be +*killed* rather than to *stop*, and an apply interrupted that way is the half-configured machine +this design is about. So it traps the shutdown signal, passes it down, and waits. + +**A clean exit is the upgrade path**, and it is the easiest thing to get wrong — twice now. The +host stands aside for a new binary by exiting zero, so anything supervising must restart on a +zero exit and must not count it as a failure. + +**Recovery is local, and detection is the mesh's.** Nothing dials a node and a host that cannot +start cannot report, so the node must recover itself. But a local supervisor sees one process +failing and cannot tell a broken machine from a broken release — only something watching every +node can, which is why a host rollout is staged and stops when nodes go quiet. + +## A host may be episodic + +**Resident or episodic, and both are hosts.** A phone has no init to register with and nothing +worth supervising, because a supervisor would be killed alongside what it supervises. So it runs +when the platform allows and is killed when the platform wants the memory — **and that is +disconnection**, which is already an ordinary situation. + +It needs no keep-alive and no new mechanism: the store is already authoritative while +disconnected, reconcile already happens on start, and *last heard from* is already reported +rather than alarmed on. An episodic host cannot be the first node, because every bootstrap step +is a shape it refuses. + +## What a declaration is + +**An ordered list of resources the host owns.** JSON, because Go's standard library carries a +JSON parser and no YAML, and the one binary whose argument is that it needs nothing must not +gain a parser to buy authoring comfort in a machine-written document. + +**Ordered, because ordering is a decision.** The host does not sort and does not resolve +dependencies — that would be deciding, and deciding the thing most likely to differ between what +the control plane intended and what the machine does. + +**Every resource has a stable identity** — a name the control plane keeps across declarations, not +a position and not a hash of content. It is what lets the store say *this is the same resource I +applied last time*, which is what makes removal possible at all. + +**Unknown is refused, whole.** A field the host does not know is something the control plane +believes it asked for. A declaration naming one is rejected entirely, naming every problem at +once — a host that applied the parts it understood would leave a machine that looks configured +and is not. + +**Six shapes:** `file`, `directory`, `service`, `package`, `container`, `action`. Every addition +widens what a compromised control plane can express, so the list is a security artefact and grows +deliberately. + +### The bundle may carry actions; the link may not + +An `action` runs a command, and the host never learns what it means. It is needed because the +bootstrap creates a database before there is any mesh to ask for one, and the host must not learn +what a database is. + +**Permitted from the bundle, refused from the link**, and the asymmetry is the whole point: a +bundle arrives *with* the binary, so anyone able to put a hostile action there could have put it +in the host itself — refusing it buys nothing and costs the bootstrap. The link is a separate +party, reachable separately, and an action there is an unbounded blast radius. + +**An action must carry its own verification**, which is also its idempotency check. The host does +not know what a database is, so *is it already there* is a question only the declaration can ask. + +## Consequences + +- **The migration is smaller than it looks.** A joining node never needs mesh-wide state — it + needs an identity, an address and one peer, and the rest arrives as declarations. +- **What is applied is recorded after it works, never before.** A failed apply leaves the machine + in whatever state it reached, and nothing must claim otherwise. +- **A second operating system is additive**: two appliers and a four-line init file. diff --git a/02-DECISIONS/0006-the-substrate-and-the-control-plane.md b/02-DECISIONS/0006-the-substrate-and-the-control-plane.md new file mode 100644 index 0000000..37f408c --- /dev/null +++ b/02-DECISIONS/0006-the-substrate-and-the-control-plane.md @@ -0,0 +1,279 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-28 +deciders: jochen +reconstructed: false +--- + +# 6. The substrate and the control plane + +*Consolidated 2026-08-28 from six records. Extended 2026-08-29, by building it: the language, and +what must be running before the control plane starts — which this record had left not established +and could not have settled the way it was asking.* + +## The control plane is what needs to know about more than one node + +That is the whole test, and it follows from the host applying rather than deciding: **deciding +needs knowledge a single machine does not have.** + +| question | whose | +|---|---| +| write this file, with this content, with this mode | the **host** | +| which nodes should run the store | the **control plane** | +| is this unit running | the **host** | +| which peers belong in this node's overlay | the **control plane** | +| has this node been unreachable for a week | the **control plane** — nobody else is watching | + +**Anything a single machine could answer alone is not the control plane's.** + +### Seven contexts and one interface + +**inventory, config, connectivity, provisioning, delivery, observability, identity** — plus +`api`, the one interface every surface speaks to. Each earns its place by the test above rather +than by being ours. + +**`work`, `knowledge` and `stream` are mesh-hosted applications, not control plane.** A task does +not need to know a node exists. *Being ours does not make something infrastructure.* + +**`identity` owns SSH access.** *Written 2026-08-29, on noticing it was assumed everywhere and +stated nowhere.* SSH appears three times across this design and every time as something that +*uses* the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary +traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only +thing that knows which humans and agents exist and which nodes they may reach, which is +`identity`'s definition. The node end already works, since an `authorized_keys` file is a file. + +It is three questions wearing one name, and only two of them are the mesh's: + +| | | +|---|---| +| **humans** | their key, on the nodes they are allowed on | +| **agents** | the same, with a lifetime — and revocation that has to actually work | +| **an agent reaching another node** | **this is the point of the mesh, not an exception to it** — see below | + +**There is no such thing as node-to-node SSH here, and that is a clarification rather than a +restriction.** The actor is always an **agent**; a node is only where it happens to be running — +[ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md): *a node is a place where an agent can +run, that is the entire relationship.* An agent hired onto one node reaching another to do work is +the capability the whole arrangement exists to provide. + +**The credential is the agent's, never the node's.** It lives in the agent's own credential +directory ([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)), so a node's +`authorized_keys` lists **agents** and never nodes. Three things follow, and they are why this +shape is better rather than merely allowed: + +- [ADR 0004](0004-a-node-and-how-it-joins.md)'s *a node holds its own identity and nothing else* + stays true — no node holds a key that reaches another node; +- a compromised node costs the credentials of the agents that were on it, not a way into + everything; +- **who may reach what is a mesh-wide fact**, which is exactly why it is `identity`'s and not + something arranged locally. + +**And it does not conflict with the host having no inbound control surface.** That rule is about +how a node's *declared state* changes: over the broker, never by being dialled. An agent with a +shell is not the mesh reconfiguring a machine — it is what a person with a terminal has always +been, and this design already depends on it working +([ADR 0007](0007-connectivity.md): the overlay is *the way back in*). What such a session can leave +behind is drift, and drift is what reconciliation is for. + +**Where the record lives is deliberately open.** Contexts integrate through it, which makes it +load-bearing, and putting it in the substrate risks recreating the circularity the tiers just +removed. Listing it as an eighth context would settle by naming what has not been settled by +arguing. + +### One node runs it, and nothing takes over + +**Declared, never elected.** No promotion, no quorum, no fencing, no split brain — none of it +built, so none of it can be subtly wrong. + +#### The option that would make it a real mesh, and why not + +*Written 2026-08-29. It had been rejected by never being written down, which is the weakest way +to reject anything.* + +A genuine peer-to-peer mesh means **no node is special**, and that has a concrete price: + +- every node holds the **whole inventory**, so there is a replication process between them; +- replication needs a writer, so one node is elected **master**, and something promotes a new one + when it drops — Redis Sentinel and its whole family of problems; +- and it still would not deliver what the name promises, because **application databases are not + replicated.** A workload's store lives where the workload lives. + +That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have +to become **a replicated database system for everything running on it** — not for our own +inventory, for every consumer's data too. That is a product, and a much larger one than the thing +it would be supporting. + +**So there are three central roles, not one**, and it is worth seeing them separately because +only the third costs operation: + +| | its loss costs | +|---|---| +| **the control plane** | nothing can be *changed*. Nothing stops running | +| **the broker** | nothing can be told anything, or report anything | +| **the hub** | nodes in different places **cannot reach each other** ([ADR 0007](0007-connectivity.md)) | + +**Whether these are one node is not decided here.** All three must be dialable by every node, which +pushes toward one; nothing says they must be. + +**That is sound rather than merely cheap**, because the design already tolerates its absence by +construction: a node reconciles from its own store and never needed to ask anybody to hold the +state it was last given. **The control plane being down is not a new failure mode — it is every +node in the ordinary disconnected situation at once.** What is lost is *change*, not *operation*. + +The honest half: this node is a single point of failure, recovery is **restore rather than +failover** — which makes backup the availability mechanism rather than hygiene — and +**certificate renewal is the clock.** An outage outlasting a renewal window expires every public +name, which turns an inconvenience into an outage on a timer. Nothing measures that today. + +## The authority is the control plane, not a database + +**There is no single mesh database.** Each context owns its store exclusively, and *the mesh +database* names a thing that will not exist. + +**No node reads any of them** — not for writes, not for reads. A node is *told* what to own, over +the link, in a bounded vocabulary; it **states** what it applied, and the owning context writes. +The difference is the security boundary: something that can write cannot be prevented from +writing anything. + +**A node runs from its own store always, not as a fallback.** The current arrangement's nastiest +property is that *a node running from cache looks identical to a node running from the database*, +with no age on the cache and nothing reporting divergence. Under this there is no second mode to +be mistaken for the first. + +**What survives from the original decision:** the repository defines what exists, the mesh defines +what runs where, and no node-to-module mapping is ever committed. That is what makes the +repositories node-agnostic and why anything about the mesh can be published at all. + +**The error underneath was a category error**: *source of truth* named a storage location when it +meant an **authority**. Once the store is the answer, *which database* becomes the question, and +shared schemas follow. + +## The substrate is what the control plane consumes and cannot grant itself + +Every module needing a database asks provisioning for one. The control plane needs a database too +and cannot ask itself, because it is not running yet. **That circularity is the definition**, and +anything on the wrong side of it is raised from the bundle the host carries. + +| role | product | | +|---|---|---| +| relational store | **PostgreSQL** | its own state lives there | +| message bus | **LavinMQ** | it cannot grant itself a virtual host — and **precedes it**, below | +| object store | **MinIO** | it cannot grant itself a bucket | +| image registry | **an OCI registry** | it cannot grant itself a repository | +| identity provider | — | **conditional**: substrate only if the control plane delegates authentication, which is undecided | + +**The role and the product are both written.** The role is what the argument turns on; the product +is what gets installed and pinned, and a design that names only the role does not record that the +choice was made. **The dependency is on the protocol** — AMQP, S3, OCI — which is what keeps +naming them safe. The store is the exception: the provisioning model uses databases, roles and +schemas as PostgreSQL means them. + +**A container runtime is detected, not chosen** — docker or podman, because a machine that +already has one keeps it. Only the version probe differs between them; the behavioural difference +(podman has no daemon, so containers do not return after a reboot unless a unit is enabled) +belongs in the declaration rather than the host. + +**Being substrate and being in the bundle are different questions.** PostgreSQL and LavinMQ must +precede the control plane; the object store and the registry are substrate by role and ordinary by +delivery, provisioned once there is a control plane to do it. + +### Why the broker precedes it too + +Written after the fact, because this record first left it *not established* and framed it as +turning on whether the control plane's own contexts talk to each other over the bus. + +**They do not** — they are one process and dispatch internally. Under that framing the broker is +provisioned like anything else and the bundle stays at one image. + +**The framing cannot answer the question.** What decides it is not how the contexts reach each +other. It is how the control plane reaches a *node* — and that is settled above: never except over +the link, and the link is AMQP ([ADR 0002](0002-nodes-communicate-over-a-broker.md), +[ADR 0004](0004-a-node-and-how-it-joins.md)). So: + +``` +the bundle raises the control plane +the control plane provisions the broker ← by telling a host to run it +telling a host happens over the link +the link is the broker +``` + +**And it is not avoided by the first node being local.** `enrol` dials the broker at the address in +its token, which is the first node's own third step. A machine that raised the mesh still joins it +the ordinary way, and that was deliberate — its specialness lasts two commands. Making it join by +some other route would buy a smaller bundle by giving up the property the design was built to have. + +The broker precedes the control plane for the same reason PostgreSQL does: **the control plane +cannot grant itself the thing it would need in order to grant it.** + +**What it costs.** Two images rather than one, against the wish above to keep the bundle at roughly +one so a person can read it — two is still readable, four would not be. And three things that are +not images, each an action the bundle declares and the host runs, the way the database already is: +a virtual host, a credential on it, and **a certificate**. That last is the awkward one: a token +pins the fingerprint a host must expect *before it sends anything*, so the broker needs a +certificate at a moment when there is no mesh to issue one and no public name to obtain one for. +Self-signed and pinned is the shape that fits; how it is later replaced by the certificates in +[ADR 0007](0007-connectivity.md) is not decided here. + +## The control plane is written in Go + +The same language as the host, so tiers 0 and 2 are one language and not two. + +The reason that decides it is not familiarity. **Its image is pinned by digest in the bundle**, +which means it is fetched and run on a machine where no mesh exists yet — nothing to check it +against, nothing watching, and a person expected to have read the bundle and believed it. A +statically linked binary makes that image the program and nothing else: no interpreter, no package +tree, no transitive dependency that arrived because something needed a date library. Everything +under that line is something somebody would have to audit, on the one image the whole mesh is +raised from. + +A second reason, smaller and still real: the control plane runs a reconcile loop of its own +([ADR 0010](0010-delivery.md) — artifacts against source, as the host reconciles machine state +against declarations). Two loops of the same shape are cheaper to hold in one head when they are +also the same language. + +**The option rejected** is TypeScript, matching the lab and the surfaces that will speak to this. +The argument for it is that tier 3 is web and CLI, so a TypeScript control plane would share types +with its callers rather than generating a contract. True, and it does not reach far enough: +`mesh-sdk` is *contracts shared across tiers* and **tier 0 is Go**, so the contracts cross a +language boundary whatever tier 2 is written in. The choice is between generating them for one +consumer or for two. + +**What it costs, plainly:** the control plane can import nothing that exists today, and a person +moving between tier 2 and tier 3 changes language. Neither is recovered later — the language is the +most expensive thing in this record to reverse. + +## The installer fetches what it pins + +`substrate.lock` carries **references, not payload** — an image name and a **digest**, fetched at +apply time. A tag moves; a digest does not, and reproducibility comes from pinning the identity of +a thing rather than carrying its bytes. + +The assumption that a machine might have no network came from the lab and was wrong: a machine +being adopted has one, and the sealed case is the lab. + +**The lab places images by raising a registry inside the scenario**, which is what a real node +pulls from anyway — so it tests the real path rather than a stand-in for it. The digests that +registry serves are its own, and that satisfies this rule: what is required is a reference that +is **exact and cannot move**, and one it assigned is both. Assuming an upstream digest had to be +preserved is what made this look impossible for a while +([04-ISSUES/009](../04-ISSUES/009-a-digest-pinned-image-cannot-be-placed-in-the-lab/00-report.md)). + +**Its contents are per operating system** even though its mechanism is not — package names, unit +names and service names all differ, so an Arch host embeds an Arch bundle. + +## Consequences + +- **The bundle stays small and reviewable.** A list of pinned references is something a person can + read; a bundle containing images is not. +- **An apply can fail because something is unreachable**, which a self-contained artifact could + not. That must fail *legibly*, naming what could not be fetched and from where. +- **Cross-context reporting is harder, and that is the point.** Anything wanting to see across + contexts consumes their events or calls their interfaces. +- **A queue with no limit grows until the broker's disk is full**, and the broker is what every + node depends on. The bound is per queue and is not decided. +- **The bundle carries two images and four actions**, and the substrate bootstrap grows a step. +- **Nothing in the first node's path is special-cased.** Enrolment is walked on node one. +- **The broker's certificate at bootstrap has no answer yet**, and is named as unfinished rather + than assumed. It is the first thing that will be wanted when the link is built. +- **The language cannot be revisited cheaply.** It is the one line here close to irreversible. diff --git a/02-DECISIONS/0007-connectivity.md b/02-DECISIONS/0007-connectivity.md new file mode 100644 index 0000000..f70d784 --- /dev/null +++ b/02-DECISIONS/0007-connectivity.md @@ -0,0 +1,191 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-28 +deciders: jochen +reconstructed: false +--- + +# 7. Connectivity + +*Consolidated 2026-08-28 from three records. Overlay, resolution, exposure, filtering and +certificates are one design.* + +## Why it is control-plane work + +Apply the test — *everything that needs to know about more than one node* — and not one of the +five can be answered by a machine on its own: + +| | needs to know | +|---|---| +| **overlay** — who peers with whom | every node, and which can be dialled | +| **resolution** — which name is which node | every node | +| **exposure** — which public name reaches which container | which node is publicly reachable | +| **filtering** — which port is open, to whom | what is assigned here, and the overlay's shape | +| **certificates** — who may present which name | which name belongs to which node | + +That is exactly what the current arrangement gets wrong, by computing all five on the node from a +direct database connection. Two modules do this, and they are the only two left holding a +credential to the control plane's database. + +**The shape of the fix, once for all five:** the connectivity context computes the configuration; +it arrives over the link as `file` resources; the service reads files and knows nothing about the +mesh. **This costs no new host vocabulary.** + +## A route is a grant + +**Ingress is not substrate.** The control plane does not need a route to start — it listens +locally — and no node needs one to reach it, because the node dials out and has no listening +control surface. It grants itself a route afterwards, the way it grants itself a bucket. + +The strongest objection deserves stating: the `api` is the one interface every surface speaks to, +so eventually it *does* want a public name. But **wanting one later is not needing one to +start**, and that distinction is the entire substrate test. + +**A module that must be reachable declares it needs a route; the proxy provides one.** Ordinary +instantiation, with the direction mirrored — the consumer supplies a target and receives a name. + +**Exposure is three facts at two scopes**, which is why it cannot live on the node: + +| the fact | scope | +|---|---| +| the public name resolves to an address | **mesh** — which node is publicly reachable | +| a certificate valid for that name exists | **mesh** — issued once, used on one node | +| the proxy maps that name to that container | **node** | + +**A node without a public address is proxied by one that has**, across the overlay. Most nodes sit +behind a connection with no forwarded port, so exposure cannot assume the workload's node is +reachable. + +## Reachability is declared, not inferred + +The overlay's peer graph is computed from whether a node can be dialled, and that was inferred +from a regular expression over the address. **The address is evidence of reachability; it is not +the fact**, and the gap has already cost: + +| address | the regex says | actually | +|---|---|---| +| `100.64.0.0/10` — carrier-grade NAT | **public** | **not reachable.** An endpoint is written to an address nothing can reach | +| any IPv6 address | public | the test is v4 shapes only | +| a routable address behind a closed firewall | public | not reachable | +| a documentation range standing in for a public segment | private | reachable — this is the lab bug | + +**A test environment having to choose its addresses to satisfy a regex is the regex telling us it +is not a fact.** + +So: **an endpoint, or none** — declared. And **the hub is declared, never derived from an address +prefix**, because an election decided by the first four characters of an address fails silently, +cannot be queried, and makes a renumbering an outage. + +**The address remains evidence and stops being the fact.** Where an observed endpoint disagrees +with a declared one, the disagreement is a **reportable condition**, not a silent correction. + +**What does not change** is the lesson underneath: role does not imply reachability — a +home-hosted node is a server that cannot be dialled. This keeps that and stops encoding it as a +pattern match. + +### Some nodes must be reachable, and this had not been said + +*Written 2026-08-29, on being asked and finding no answer.* + +Everything above treats reachability as a **fact to record** — which node can be dialled, so that +exposure and certificate issuance can be placed. It never said the converse, and the converse is a +hard requirement: + +| role | dialled by | so it needs | +|---|---|---| +| **the node running the broker** | every node, outbound ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) | to be reachable from wherever nodes are, at a **stable address** | +| **the hub** | every node not co-located with its peer | the same | + +**Across the internet, "reachable from wherever nodes are" means publicly reachable.** For a mesh +confined to one network it does not — the requirement is about the nodes that exist, not about the +public internet. + +**Stable is the sharper half.** A token carries the broker's *address, not a name*, because there +is no resolution before joining ([ADR 0004](0004-a-node-and-how-it-joins.md)). A broker node whose +address moves invalidates every token issued for it, and a node that was disconnected across the +change cannot get back. + +**A mesh whose nodes are all behind NAT cannot be raised.** That is a real precondition and it +belongs with the others rather than being discovered. + +### The link stays on the underlay, and that is a repair channel + +The obvious objection is that all traffic should run over the overlay. Nearly all of it does — SSH, +services, node to node — and the exception is each node's own outbound link to the broker. + +**At join time it is forced**: a node has no overlay yet, so it cannot use one to ask for one. +**Afterwards it is a choice**, and the reason is that the link is how a broken node is fixed. A +repair channel carried over the thing being repaired is not a repair channel: a node whose only +path home was the overlay is gone the moment an overlay declaration is wrong. + +**What it does not cost is confidentiality.** The link is already authenticated and encrypted +against a pinned fingerprint ([ADR 0004](0004-a-node-and-how-it-joins.md)), so moving it onto the +overlay would not protect traffic that is unprotected today. + +## A filter rule names its source + +`scope: public` is declared in five manifests, is part of no rule type, and is **referenced by no +code**. So five manifests appear to restrict a port and restrict nothing — on the modules most +worth restricting. + +**A rule names its source. `from:` is the only way to scope one, and a rule without one is open** +— which it must say plainly rather than appear to deny. + +**`scope:` is removed rather than implemented**, because giving it meaning would leave two ways to +express one thing. And the general fix is that **an unknown key is refused**: the host's +declaration parser already works this way, and manifests are the layer where that discipline is +missing. `scope:` survived because nothing rejected it, and it spread by copying to five +manifests. + +## Order, and what it costs + +**The link runs on the underlay and never on the overlay.** The overlay is configured by the mesh, +so a link requiring it could never be established on a new node. + +**The first declaration is the overlay and nothing else** — because a node's address and peers are +*assigned* so it cannot come earlier, and because it is the way back in. A node reachable over the +overlay can be fixed by hand if a later declaration breaks the machine; **a large first +declaration risks a node that is broken and unreachable at once.** + +**Reachable is not the same as having a control surface.** Every node reaches every other over the +overlay — SSH, services, ordinary traffic — and every node consumes from the broker. What is +forbidden is a listening thing that accepts instructions and changes the machine. + +## Consequences + +- **The last two direct database connections leave the nodes**, and with them the database + credential every node carries. +- **The `/etc/hosts` floor goes**, along with the bootstrap circularity it patched. +- **Two certificate authorities stay separate on purpose**: a public one for public names, the + mesh's own for internal ones. A single-CA lab would hide any bug living in the split. +- **What happens when the hub is down**: nothing takes over. Non-co-located paths stop; co-located + peers and every assigned workload keep running. + +## Open — the link over the overlay, with a fallback + +*Raised 2026-08-29 and deliberately left open, because the honest gain is smaller than it looks +and it is a decision rather than a derivation.* + +The proposal: a node prefers the overlay for its link and drops to the underlay when the overlay +is not working — so ordinary operation is private and the underlay stays as the way back. + +**Two things it would have to get right:** + +- **The trigger cannot be "is the overlay up".** A WireGuard interface has no link state; once + configured it is up whether or not the far end exists. So there is no flag to read, and failing + over means *try, fail, time out, retry elsewhere*. +- **Running on the fallback has to be visible.** A node that quietly drops to the underlay is a + node whose overlay is broken with nothing to say so, and it will stay broken because everything + still works. That is this repository's recurring fault — a failure that reads as success — and a + fallback is the easiest place in the design to reintroduce it. + +**What stops it being an obvious win:** if the underlay path must stay available for the fallback, +the broker stays exposed on it. So the exposure is unchanged and the traffic was already encrypted +— the gain is which network carries bytes, not what an attacker can reach. + +**The version that would buy something is overlay-only**, with the broker firewalled to the overlay +in steady state, accepting that a node whose overlay breaks needs hands-on recovery. That is a real +trade: it exchanges the automatic way back in for a closed port. + +Not decided either way here. diff --git a/02-DECISIONS/0008-a-context-owns-its-store.md b/02-DECISIONS/0008-a-context-owns-its-store.md new file mode 100644 index 0000000..8175dd4 --- /dev/null +++ b/02-DECISIONS/0008-a-context-owns-its-store.md @@ -0,0 +1,114 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-26 +deciders: jochen +reconstructed: false +extends: 0009-modules-and-the-graph.md +--- + +# 8. A context owns its store, exclusively + +## Context + +[ADR 0009](0009-modules-and-the-graph.md) settles what a module +declares. This settles what a grant may be, and it is the half that **removes** things. + +`how-we-build` §4 already says *contexts integrate through the record, never through a shared +schema*, and states the cost: several domains share one forty-five-table schema, which is why +work belonging to one context keeps having to be implemented in another. + +That was written as a principle. Counted, it is thirteen foreign tables belonging to three +separate contexts, living in the mesh's own registry database. + +## Considered options + +1. **A schema per consumer inside a shared database.** Namespaced, revocable by dropping the + schema, with a cross-context join possible but deliberate. Rejected: it keeps the letter of + §4 and leaves the temptation in place, and a boundary that is merely inconvenient to cross + gets crossed. +2. **Read-only roles on another context's store.** Rejected for the same reason and one worse: + reading another context's tables couples you to its layout exactly as firmly as writing them, + and the coupling is invisible until the owner changes a column. +3. **Exclusive ownership.** Chosen. + +## Decision + +> **A context is granted only what it exclusively owns.** + +No shared writes. No read-only role on another context's store. If you need what another context +holds, you ask it or you subscribe to it. + +**The unit is the context, not the process.** Everything inside a context — its service, its +surface, its tools — reads its own store freely. A board showing the mesh's own nodes and +modules is the mesh showing its own data, not a boundary crossing. What is forbidden is a +*different* context reading it. + +### Asking or subscribing is derived, not chosen + +[ADR 0004](0004-a-node-and-how-it-joins.md) makes disconnection an ordinary situation. So: + +- **Anything that must keep working while disconnected cannot ask** — there is nobody to ask. It + keeps a local copy, which means subscribing. +- **Anything where a stale answer is worse than none cannot subscribe.** A display may lag; a + decision about whether a grant is still valid may not. + +Neither is a query against another store, whatever transport it travels over. + +### What that is, concretely + +*Written 2026-08-29, on building the first one — the rule above was clear and what to type was not.* + +**One PostgreSQL database per context, named for the context.** A separate database rather than a +separate schema is the whole point: a cross-schema join is a qualified name away, and a +cross-database join needs a foreign data wrapper somebody has to install and explain. + +**And one credential per context, held only by it.** There is no mesh-wide connection setting and +no way to ask for one, so reaching another context's store is not a matter of restraint — a process +has no address for it and nothing to present. That is also how this rule is *checked*: what a +context can reach is the list of variables the declaration running it grants, and it is read there +rather than audited in code. + +**Contexts that do not exist yet do not get a database.** The bootstrap creates the ones there are. + +**How a context added later gets its database is open**, and it is a real question: by then there is +a control plane, but a control plane holding a credential that can create databases is holding +rather more than the thing it exclusively owns. + +## What this removes + +The first clear list of what the design deletes rather than adds: + +- **Grant kinds.** There is one: an exclusive resource. No schema grants, no read roles, no + rules about who may see what inside a shared thing. +- **The question of who owns which table**, and the guessing at revocation time. Removing a + consumer drops what it was granted, whole. +- **Cross-context migration ordering.** Two contexts migrating one database must be ordered + against each other. Exclusive ownership means a context's migrations are ordered only against + itself. +- **A class of permission modelling** a shared store would otherwise need. + +## Consequences + +- **Cross-context reporting is harder, and that is the point.** Anything wanting to see across + contexts consumes their events or calls their interfaces. That is §4's argument, and the cost + it names is the one already paid. +- **A single surface over several contexts still works** — that is what a surface is. It reads + interfaces, not stores. This holds while the contexts sit behind **one** interface; splitting + a context into its own deployable costs that, and the composition would have nowhere to live + that tier 3 permits. **A real constraint on how far the control plane may be split.** +- **Three contexts must move out of the registry database**, taking thirteen tables with them. + Their dependency on the registry then shrinks to almost nothing — one of them needs a single + table. +- **The node appliers were already handled.** [ADR 0005](0005-the-node-host.md) + stopped the host querying the mesh database for tier reasons unrelated to this, and it removes + most of the remaining direct readers as a side effect. +- **What a consumer does about events missed while disconnected is not decided** — replay from a + point, ask once and resume, or rebuild. The question every projection has. + +## References + +- [`how-we-build.md`](../00-META/how-we-build.md) §4 — the rule this makes enforceable. +- [Research 011](../01-RESEARCH/011-the-module-graph/worked-provider.md) — the count, the worked + provider, and the dashboard case. +- [ADR 0004](0004-a-node-and-how-it-joins.md) — why asking or subscribing is derived. diff --git a/02-DECISIONS/0008-a-failed-step-fails-the-job.md b/02-DECISIONS/0008-a-failed-step-fails-the-job.md deleted file mode 100644 index 7f5b006..0000000 --- a/02-DECISIONS/0008-a-failed-step-fails-the-job.md +++ /dev/null @@ -1,71 +0,0 @@ ---- -status: accepted -date: 2026-06-05 -deciders: jochen -reconstructed: true ---- - -# 8. A step that fails must fail the job - -> Reconstructed after the fact from the evidence cited below. - -## Context - -The mesh's expensive faults are not crashes. They are the operations that reported success and -did nothing: an artifact that partially downloaded and was extracted anyway, a package that -404ed from every mirror while the job went green, a hook that never ran because it was named -for a feature the module does not declare, a deploy that reported the transport succeeded -rather than that the effect happened. - -Each of these was found long after it happened, by someone investigating an unrelated symptom. -The cost is not the failure; it is the interval between the failure and anyone learning of it, -during which decisions are made on the assumption that the thing worked. - -## Considered options - -1. **Continue on error and report at the end.** Rejected — it is largely what existed. A - summary nobody reads is not a report, and later steps run against the state the failed step - should have produced. -2. **Continue on error, and let health checks catch the divergence.** Rejected. It converts a - precise, located failure into a vague one discovered elsewhere, and requires a health check - for every possible partial state. -3. **Fail the step, fail the job, say which step.** Chosen. - -## Decision - -A step that fails stops the sequence it is part of, and the failure is surfaced where the work -was requested — not only in a log. - -Concretely, and these are the forms it takes: - -- A scripted sequence gates each step on the previous one. A directory change that fails must - stop the commands that assumed it. -- An artifact that does not fully download is not extracted. -- A stage reports the **effect** it achieved, not that it dispatched a message. "Started" must - mean the thing is running, not that a command returned. -- A template that cannot resolve a variable is not written half-rendered. - -**Prefer failing to lying.** A green result that is not true costs more than a red one. - -## Consequences - -- Failures are noisier and land earlier, on the person who caused them. -- Some jobs that used to complete now stop. In every case examined so far, that job was - producing a partial result that something downstream trusted. -- This is a rule the mesh has adopted repeatedly rather than once, because each instance is - written in a different place — a shell hook, a download path, a deploy stage. It is not - enforced by a mechanism, and cannot currently be checked in general. New instances are still - being found; the package-install case remains open as - [`04-ISSUES/001`](../04-ISSUES/001-failed-package-install-reports-success/00-report.md). - -## References - -- `fix(installer): fail loudly when feature artifact download fails` (#244), 2026-06-05. -- `A flavor template with an unresolved variable is written to disk instead of failing` - (#710), 2026-08-08. -- Knowledge base: `troubleshooting/deploy-reports-transport-not-effect`, - `troubleshooting/service-started-is-not-ready`, - `troubleshooting/green-pipeline-means-transport-not-effect`, - `troubleshooting/silent-failures-and-stale-state`. -- The core value it became: [`00-META/mission.md`](../00-META/mission.md), "Failure must - be loud." diff --git a/02-DECISIONS/0009-modules-and-the-graph.md b/02-DECISIONS/0009-modules-and-the-graph.md new file mode 100644 index 0000000..08fe2ab --- /dev/null +++ b/02-DECISIONS/0009-modules-and-the-graph.md @@ -0,0 +1,441 @@ +--- +topic: what runs on it +status: accepted +date: 2026-08-28 +deciders: jochen +reconstructed: false +--- + +# 9. Modules and the graph + +*Consolidated 2026-08-28 from six records.* + +## Everything is a module + +One kind of thing, one manifest describing all of them. A database, a web application, a window +manager and a firewall rule set are all modules — not because they are alike, but because +**anything else means a second kind of thing with its own rules, and then a third.** + +**A module is the unit of delivery**: assignable to a node, versionable, replaceable on its own. + +## There are no domain modules + +An earlier decision grouped modules by domain — four things constituting *how a node is +reachable* becoming one `networking` module. **That was wrong, and the correction is worth +keeping** because the observation behind it was right. + +The measurement holds: reachability is the **only** place in the catalogue where modules +genuinely change together under one intent. What did not hold is the conclusion. Tight coupling +means they share an **authority** — one place that decides for all of them — and not that they +should be one artifact. `wireguard` and the proxy are deployed to different sets of nodes, so a +module containing both would be assigned where half of it is unwanted. + +> **Coherence is a context. Delivery is a module.** + +**Folders assert relationships; edges record them.** What grouping was for — finding things, +seeing what belongs together — is a tag and a query, neither of which anybody has to keep true by +hand. + +### What a domain module turns out to be, and why it is not the one refused above + +*Written 2026-08-29, from building it. The heading above reads as a contradiction of what now +exists and is not one — but only if the difference is stated, so it is stated here.* + +**What was refused contains things. What exists contains nothing.** + +| | `networking` as refused | `networking` as built | +|---|---|---| +| what is in it | WireGuard, a proxy, a firewall — artifacts | nothing at all | +| what it says | *these ship together* | *I want a private network and names* | +| what is assigned | one module, half of it unwanted | whatever answers each requirement, each on its own | + +The objection above is untouched by this and still correct: a module holding WireGuard and a +proxy is assigned where half of it is unwanted. **A module holding nothing cannot be, because +there is no half.** It is requirements and a name, and every artifact it leads to is still an +ordinary module assigned on its own terms. + +**Why it is worth having.** Most people want the network working and do not want to choose a VPN. +`assign networking` finds one answer to each requirement and takes it without asking, because +with one candidate there was never a question — the rule below about refusing does the work. +Somebody who does care assigns the VPN they want, and *that is the whole of choosing*: there is no +flavor field, no variant syntax, and no second verb. **Picking an implementation is assigning a +module.** + +**What it costs, stated because it is real.** Adding a second implementation to the catalogue +turns a settled question into an open one for **everyone using the bundle**, not only for whoever +wanted the alternative. Every node assigned `networking` refuses until somebody says which. That +is [the refusing rule](#a-requirement-with-several-answers-is-refused-never-guessed) applied +consistently, and the alternative is a default — which is the flavor field returning under a +better name. The cost is one assignment per node, and the message names the candidates. + +**A consequence that had to be found by running it.** A bundle can drag an implementation in +through a requirement nobody looked at. Choosing a different VPN still installed WireGuard, +because the names module needed addresses only WireGuard hands out, and nobody was told. Two VPNs +on one machine is not always wrong — a machine may run one for another purpose — but being **the** +network the mesh runs over is singular, so that is a claim, and the collision is refused by name. +**The general rule: what a bundle pulls in is only as safe as the claims on what it pulls in +from.** + +## Three edges + +| edge | means | declared? | satisfied | +|---|---|---|---| +| **presence** | that thing must exist and be reachable here | yes | at provisioning | +| **instantiation** | that thing makes something for me and hands back credentials — a database, a bucket, a route | yes | at provisioning, and again whenever it must be | +| **build** | I was compiled against that artifact | **no — read from imports** | **at build, once** | + +**Instantiation implies presence; presence does not imply instantiation.** + +**A route is an instantiation edge**, and it is worth noticing because the direction is the mirror +of a database: the consumer supplies a target and receives a *name*, rather than supplying nothing +and receiving credentials. Same edge. + +**Provider stops being a category.** Any hosted thing can be a factory — an identity provider +grants clients, a mail server grants mailboxes. It is a facet, not a kind. + +**A module may also declare what it claims**, because some things cannot coexist and that is a +fact about the module rather than about a particular node. What that means precisely is below. + +### Where the answer to a requirement is allowed to live + +*Written 2026-08-29, from building it. The table above distinguishes **presence** from +**instantiation** and this is the half of that distinction nobody had noticed was missing: not +what the edge hands over, but **where the thing on the other end is.*** + +Two different things were both being written as a requirement: + +| | *a shell*, *a display server*, *a private network* | *a database*, *an object store*, *an identity provider* | +|---|---|---| +| where the answer lives | **this machine** | **somewhere in the mesh** | +| how it is answered | install another module here | find the node already running it | +| what is missing if absent | a module to assign here | **a decision about where**, which is nobody's to make silently | + +Answering the second like the first installs a database on every machine that uses one, which is +what it did. + +**So a provided name carries a scope**, the same idea a claim already has, and written short in +the ordinary case so the few that are not node-scoped stand out rather than drowning. Scope is a +property of **the name, not of each provider**: two modules disagreeing about whether a database +is local would make one requirement mean two things depending on which happened to answer it, so +that is refused. + +**A requirement answered from the mesh is never satisfied by installing it here.** Nothing, and +the mesh refuses and says which module to assign somewhere. Two, and it refuses and says how to +choose — the same rule as everywhere else, for the same reason: picking is guessing, and the wrong +guess puts somebody's data on a machine they did not choose. + +**Choosing is recorded per node**, because that is the granularity the choice actually has — two +machines may reasonably use two different databases and a mesh-wide answer could not say so. A +choice pointing at a machine that does not provide the thing is refused rather than quietly +replaced by one that does, and a single available provider does not override a choice either. +**Both are the same rule: the mesh does not overrule a person, and it does not move data without +being told to.** + +**What this is a prerequisite for.** Knowing *which node* answers is the first half of handing a +credential back — you cannot be given a database's password before it is settled whose database it +is. So a node's resolution now records what it takes from elsewhere, which is both the only part +of its set that stops working when a *different* machine goes away, and the place a credential +will hang. + +### An edge has two directions, and only one of them is built + +*Written 2026-08-29, from building it. The row above already says a consumer **supplies a target +and receives a name**; what it did not say is that those are two separate mechanisms, and that +having one without the other is what forced two modules outside the system entirely.* + +| direction | the consumer says | who needs it | +|---|---|---| +| **contribution** | *publish me at this name, on this port* | the proxy, the DNS server, a firewall | +| **binding** | *and give me back a credential to it* | the database, the object store, the identity provider | + +**Contribution is built.** A module declares what it contributes to a requirement; the control +plane collects every contribution on a node and writes them to a path the provider named, as a +file, in the mesh's own shape. **Contributing to something is requiring it** — asking to be +published means a publisher must exist, and a module that had to say both would eventually say +one, with the failure appearing as a machine where nothing serves the route. + +**The control plane does not know what a reverse proxy is**, and does not write one's +configuration. It delivers the facts; the module turns them into whatever it runs. That boundary +is what makes swapping the proxy cost nothing in any module that publishes through it, and it is +[the same separation](0001-mesh-brokers-nodes-host-agents-think.md) that keeps third-party +software *on* the mesh rather than *of* it. It also costs the host nothing: a received file is a +file, which was checked by putting the control plane's output through the host's own parser rather +than by asserting it. + +**Binding is built except for the secret**, and that turned out to be the useful way to cut it. + +A provider says what a consumer needs in order to use it — a port, a driver, a realm — and a +consumer says where it wants to be told. The mesh adds the half only it has: **which machine, and +what that machine is called on the private network.** So an application on one node is handed the +address of its database on another, as a file, and reaches it by a name the mesh also created. + +**The file states that it carries no credential, and why.** A missing field looks like a bug; a +stated absence looks like a boundary, and somebody wiring this up should not spend an afternoon +looking for a password that was never going to be there. + +### And the secret, which is delivered without ever being held + +*Written 2026-08-30, after looking at how the existing mesh does it. The design here is a reaction +to a measurement, not a preference.* + +**The obvious arrangement is a credentials column, encrypted at rest.** It exists, and its own +tooling records what it bought: + +| | | +|---|---| +| the tool for finding a secret matches **by value**, not by name | because one password is in the provisions table, the environment table, each node's environment file in plain text, and **inside every connection string composed from it** — copies its documentation calls *"often the only copies actually in use"* | +| a query against the encrypted column **returns zero rows and proves nothing** | so auditing moved to the decrypted copies on the machines | + +**Two faults, and encryption at rest addresses neither.** The control plane can read what it +stores, so a copy of its database is a copy of every credential in the mesh. And one secret has +many homes with nothing tracking them — **composition is what mints the untracked ones**, because +building a connection string centrally creates a new secret-bearing value no rotation path knows +about. + +**So the value is sealed to the node that will use it before it is stored.** With a key that node +generated and whose private half the mesh has never seen — a third key beside the identity and the +overlay, for the same reason those are two rather than one. What is stored is unusable by whoever +holds it, the mesh included, and the broker relays a blob it cannot read. This is what makes +[ADR 0004](0004-a-node-and-how-it-joins.md)'s *compromise of a node is compromise of that node* +true of secrets rather than true of identity and quietly false of everything that matters. + +**And nothing is composed centrally.** A connection string is assembled on the machine that needs +one, if at all. The mesh delivers parts. + +**What it costs, stated because it is real:** the mesh cannot audit by value. That is the right +trade rather than an oversight — a query over an encrypted column could not either, so the audit +was never real. What *is* answerable is which node holds what, which is the question rotation +actually asks. + +**A consequence that shapes the mechanism.** The mesh discarded the plaintext, so it cannot +compose a file containing it. The credential is therefore **its own file**, holding the value and +nothing else, beside the readable one. That is better than the alternative it was forced into: +the readable half stays readable in the declaration, and the secret half changes only when the +secret does, so a service reloading on it reloads for a real reason. + +**Rotation is generating a new one**, because reading the old one back is not possible. Both ends +are re-sealed and reach their machines in the same push — which removes the window where half the +mesh holds a dead credential, the failure +[recorded in ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) as consumers on three nodes +holding one for two days. + +### The provisioner, which is where the mesh stops + +*Written 2026-08-30, from building one and running it against a real database.* + +**A password nothing was told to create authenticates nowhere.** The mesh generates one, seals it +to both ends and cannot read it — so it cannot tell the software to start accepting it either. +Something on the providing machine reads what arrived and makes it true. That is a provisioner. + +**It belongs to the module, not to the mesh**, and the boundary is the same one that keeps +third-party software running *on* the mesh rather than being *of* it +([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)). The control plane decides and never +touches a machine. **What the mesh owns is the contract**, which is two files the host writes from +an ordinary declaration: + +| | | +|---|---| +| the manifest | every consumer, what it asked for, and **where** its credential is | +| one file per consumer | that credential, alone in it | + +Two files because the mesh discarded the value and cannot compose a document containing it. As +before, the constraint produces the better shape: the readable half stays readable and auditable +in the declaration, and the secret half changes only when the secret does. + +**It reconciles; it is never told what changed.** It runs after every declaration and must reach +the same state from wherever it starts. Three consequences, and each of them is a fault that has +been shipped somewhere: + +- **the password is set every time, not only on creation** — otherwise the role already exists, + nothing happens, and a rotation reports success while changing nothing +- **what it made and nobody asks for any more is removed** — otherwise a consumer that left keeps + a working login for ever and nothing ever says so. This is the same rule the host follows about + [removing what it declared and no longer declares](../04-ISSUES/010-the-first-declaration-destroys-the-substrate/00-report.md) +- **what it did not make is left alone** — otherwise it cannot be run on a system that predates + it, which is every system anybody would want to adopt + +**A missing credential is refused rather than worked around.** A role created without one is a +login nothing can use, and nothing would report it until something tried to connect. + +**This is where the mesh stops**, and saying so is the point of the section. It decides, delivers +and can prove what it delivered; the last inch belongs to whoever knows what `create role` means. + +**One check that only became possible now.** Two machines wired together across no private network +is a mesh that reports itself configured and does not work, and the failure surfaces as a +connection timing out — the slowest place to find anything. It is refused, and it is only +*checkable* because the network became [something a machine is +given](#what-a-domain-module-turns-out-to-be-and-why-it-is-not-the-one-refused-above) rather than +something it has by virtue of holding an address. + +**What the absence cost, measured.** Exactly two modules opened a direct connection to the control +plane's database — the proxy and the VPN — and they are the reason every node permanently holds a +credential to it. Both were doing by hand what this edge is for. The VPN's half is closed by being +[a module whose files are computed](../03-DESIGN/01-to-be/08-connectivity.md); the proxy's is +closed by contribution. **Neither needed a new kind of thing, and both had been outside the model +for as long as there was one.** + +### Why the build edge is a different kind + +It is fixed inside an artifact rather than negotiated when something runs, and **its only remedy +is a rebuild** — nothing can re-provision it. + +It is also **derived rather than declared**, and the asymmetry is deliberate: a runtime edge is an +*intention* somebody has about how the mesh should be wired, and only a person can state it. A +build edge is a *fact about code that already exists*, and a declared list of dependencies drifts +from the imports it describes. + +**An artifact is out of date when its source moved, or when anything it was built against moved.** +So what is recorded is a commit *and the identity of every artifact it was built against*, which +is what makes the rebuild set computable and *is this current?* answerable without building. + +**The graph measures design quality, not just build order.** A module with many inbound build +edges is one whose every change is expensive — and that is readable before anything is built. The +current shared library is exactly that, and nobody could see it because nothing drew the edges. + +## Provisioning is declared, never configured by hand + +A module declares what it **provides** and what it **requires**. The mesh satisfies it: a +provisioner belonging to the provider creates the resource and its credential, records the grant, +and the values are derived onto the consumer. **Neither the credential nor the topology is ever +written by hand.** A requirement may name a provider on another node, so cross-node wiring is the +same declaration. + +## What a module claims, and why it is not a list of rivals + +*Written 2026-08-29, replacing pairwise exclusion.* + +**Exclusivity is not a property of a module. It is a property of a singular resource the module +takes over.** Two shells do not compete for anything and any number may be installed. Two display +servers both want the seat, and only one may have it. + +> **A module declares what it *claims*. Two modules claiming the same thing cannot both be +> assigned within that claim's scope.** + +**Not "xorg conflicts with wayland".** Pairwise exclusion has a property that only shows up later: +adding a third display server means **editing xorg and wayland to know about it**. Every new +module requires changing modules nobody who wrote it owns, and the edits grow as the square of +the count. With a claim, the third one says `claims: the seat` and nothing else changes anywhere. +**The new module is the only thing that has to know anything** — which is the difference between +a catalogue that grows and one that calcifies. + +The pattern is common enough to be worth listing, because seeing it is most of understanding it: + +| these coexist | these claim one thing | +|---|---| +| shells — bash, zsh, fish | display servers — xorg, wayland (*the seat*) | +| editors — vim, emacs, helix | init — systemd, openrc (*pid 1*) | +| language runtimes | container runtime — docker, podman | +| terminal emulators | reverse proxies — nginx, caddy, traefik (*ports 80/443*) | +| browsers | time — chrony, timesyncd, ntpd (*the clock*) | +| | resolvers — resolved, dnsmasq, unbound (*`/etc/resolv.conf`*) | +| | network management — NetworkManager, networkd, netctl | +| | mail — postfix, exim, msmtp (*port 25*) | +| | audio — pipewire, pulseaudio (*the device*) | + +**A claim has a scope**, because not everything singular is singular per machine: + +| scope | example | +|---|---| +| **node** | the seat, pid 1, port 443 | +| **site** | a DHCP server on a segment | +| **mesh** | the hub, the control plane | + +The last is not new — the mesh already enforces exactly one hub with a unique index +([ADR 0007](0007-connectivity.md)). Scope is that idea, said once rather than hard-coded per case. + +**Some conflicts need no claim at all.** Two modules declaring the same file, or binding the same +port, are visible from *what they declare* — the mesh already holds every resource of every +declaration. So a claim is only written for the abstract ones, where nothing in the declaration +reveals the clash. That keeps the manifest small, which is worth protecting. + +## A requirement with several answers is refused, never guessed + +A module requiring *a shell* may be satisfied by three. The mesh does not pick. + +| candidates | what happens | +|---|---| +| exactly one | assigned, silently — there was no choice to make | +| none | refused, naming what is missing | +| several | **refused, naming them**, and a person chooses | + +**This is what makes a solver unnecessary.** Counting candidates is a few lines and has no +surprising behaviour; a solver that picks has to be understood before its answer can be trusted, +and it is understood by whoever is debugging it at the time. Nothing here is lost by waiting — +a solver can be added later without changing a single manifest, and the reverse is not true. + +**Requiring a module and requiring a capability are different fields**, because the remedies +differ and the message should say which: + +- *i3 needs xorg, which is not assigned here* — assign it. +- *this machine has no seat* — wrong machine; nothing can be installed to fix it. + +### A capability may carry a value, and that is not a new idea + +A capability is a named fact about a machine, **detected and never assumed**. Its presence gates +an assignment; its detail can also carry a value — `seat: card1-DP-1`, `panel: oled`, an +architecture, an amount of memory. Nothing new is needed for that: a verdict has always had a +detail beside its yes or no. + +So *can this run here* and *what should it be configured as* are answered by the same fact, read +two ways. A module that must not be assigned without an OLED panel and one that dims itself +differently on one are reading the same line. + +**What keeps the set from sprawling is the cost of adding one.** A capability must be detected, +and the detector must say how it knows — so nobody can add one they cannot check, which is the +whole of [04-ISSUES/007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md): +an installed package was treated as a capability and a node was assigned work it could not do. + +**And detectors ship inside the host**, which is one statically linked binary. Adding a capability +means shipping a new host to every node that needs it. That is a real cost and it argues for +keeping the vocabulary small and general — `seat`, not `has-nvidia-with-two-outputs`. + +## "Flavor" is retired + +It was carrying three unrelated meanings — variants of a thing, a subset of one module a node +installs, and whatever the current system does, which earned two knowledge-base entries about +going wrong. **A word with three meanings cannot be reasoned about**, and every attempt to design +around it produced a rule that was right for one meaning and wrong for the others. + +What it was reaching for is two ordinary things: + +- **Different modules that provide the same thing.** `zsh` and `fish` both provide *a shell*. They + are two modules, not one module with a switch: they share a name and nothing else — different + packages, different configuration, different everything. +- **One module with a setting.** A monitoring module that is an agent here and a server there is + one module, configured. Nothing varies but a value. + +If something is neither, it is probably two modules. + +**A third thing it was reaching for, added 2026-08-29:** *I want this working and I do not care +which one.* That is a module with requirements and no files — +[a domain module](#what-a-domain-module-turns-out-to-be-and-why-it-is-not-the-one-refused-above) — +and it is what makes "different modules that provide the same thing" bearable for somebody who +does not want to know there is a choice. + +## The core library is the mesh's domain + +One module everything may depend on. It holds **what is true of the mesh regardless of which +context you are in**: a module, a node, an assignment. + +The test: *would this still mean the same thing in a context that had never heard of the one it +came from?* A node would. A pipeline stage would not — that is delivery's. + +**Types ship with the module that owns them**, not here. A consumer needing `inventory`'s types +depends on `inventory` — one narrow, visible edge — rather than everything depending on a hub +where the relationship cannot be seen. **A library everything depends on is expensive to change +whether it holds types or code; the fan-in is what makes it expensive**, which is why *types, not +behaviour* was the wrong guard. + +**It stays small on its own.** A domain model changes when what the mesh *is* changes, which is +rare. A drawer labelled *shared* changes whenever anybody writes something reusable, which is +constantly — and *who else might want this* always answers yes, which is how the current one grew. + +## Consequences + +- **Fewer things will be shared, and some code will be written twice.** That is the trade: the + current library exists because sharing felt free. Two similar functions in two modules is often + the better answer. +- **The check is a measurement rather than a prohibition.** Inbound build edges say when something + is becoming a hub, while it is happening rather than after. +- **Reading build edges needs a language-aware tool per language**, which is the real cost and the + reason declaring them looks tempting. It is still wrong. diff --git a/02-DECISIONS/0010-delivery.md b/02-DECISIONS/0010-delivery.md new file mode 100644 index 0000000..adf3f97 --- /dev/null +++ b/02-DECISIONS/0010-delivery.md @@ -0,0 +1,161 @@ +--- +topic: what runs on it +status: accepted +date: 2026-08-28 +deciders: jochen +reconstructed: false +--- + +# 10. Delivery + +*Consolidated 2026-08-28 from five records.* + +## Delivery is a comparison, not a pipeline + +**The control plane holds what source exists and what has been built from it, and builds the +difference.** A change becomes a build because source is **ahead of artifacts** — answerable at +any moment — rather than because a message arrived. + +**An event makes it fast. Nothing makes it necessary.** A missed notification costs latency and +cannot cost correctness. + +That is the same shape the host uses on a machine, one layer up: + +| | reconciles | against | +|---|---|---| +| the control plane | artifacts | source | +| the host | machine state | declarations | + +**This is not the current coordinator repaired.** That is a state machine over stages; the value +of it here is as a catalogue of the ways this fails, and it has been used for exactly that. + +### The module system is the CI/CD + +*Written 2026-08-29, because this was the intention throughout and was never stated in one line.* + +**There is no pipeline product beside the mesh, and there is not going to be one.** A module +declares what it is ([ADR 0009](0009-modules-and-the-graph.md)); the control plane notices its +source is ahead of its artifacts and builds it; the graph says what else that invalidates; the +node that should run it is told. Build, test, publish and deploy are the same reconciliation seen +at four points, not four stages wired together. + +**Which is why the module system is the core of the setup rather than one component of it.** Every +other layer is carried by it: the substrate is modules the bundle raises before there is a mesh, +the control plane is a module, and an application is a module with a different manifest. A thing +that cannot be expressed as a module cannot be delivered at all — that is a real constraint, and it +is the one keeping a second delivery mechanism from growing beside this one. + +**What disappears is the pipeline as a state machine** — no stage list something can be omitted +from, which is how a verify stage was built and never scheduled, and no run to lose. + +### Currency is the whole input closure + +**An artifact is out of date when its source moved, or anything it was built against moved.** So a +shared library changing invalidates everything with a transitive build edge to it, in dependency +order, because a module cannot be built against a new library until it exists. + +**The module graph is a prerequisite of this, not an enabler of it.** Without it there is no +rebuild set and no ordering, and this cannot be implemented. + +## An artifact is build output, never a source tree + +Compiled and bundled with its dependency graph inlined. **A deploy is extract-and-run and touches +no network.** + +The consequence is the whole cost of the decision: **anything not in the build output does not +ship.** Every file kind had to be brought into that rule separately, and each was discovered by +something silently not happening after a deploy — migrations reading a source layout, +provisioning scripts reading a source layout, selection files never packaged at all. + +## Three silos, and the third is not a stage + +The cardinality observation holds and is what the split is for: + +| silo | runs | ends with | +|---|---|---| +| **build** | once per module | a self-contained artifact | +| **publish** | once per module | that artifact addressable — an image by digest, a package in the mesh's repository | +| **deploy** | **once, not once per node** | the affected nodes' **declarations updated** | + +**Deploy stops sending commands to nodes.** It changes what the control plane says each node +should be, which is one write. What happens on the machines is the host's ordinary reconcile. + +**Why this fixes the failure class rather than patching it.** Every recorded fault shares one +shape: *the thing that reported success was not the thing that did the work.* A coordinator +dispatching a command can only report on dispatch. Under this the reporter **is** the applier — +which already refuses to record a resource until it read it back, and already fails the whole +apply on one failed step. + +**The verify stage disappears as a stage**, which is the strongest evidence for the shape: +verification stops being a step that can be omitted from a list and becomes a property of applying +at all. + +**There is no fan-out**, so the defect class that came from the build node having passed through +two silos while others had not cannot arise. + +## A step that fails must fail the job + +A step that fails and lets the job continue **reports success for work that did not happen**. +Absence of an error is not evidence of an effect. + +This is the mesh's most consistent failure shape, and it is not incidental — it is what stage +reporting measured. Documented instances: a service reported started when the container command +merely returned; an image pull failure that did not fail the deploy; a package install that 404'd +from every mirror while the job went green; a node left on old code after a failed download with +a version marker that had already advanced. + +## The verdict is tiered + +An artifact may not be declared until something has judged it fit. **Two tiers, because one gate +would be both slow and unreliable:** + +| | judged by | when | +|---|---|---| +| the module's own tests | the build | **always** — this is most of it | +| the lab | a raised scenario | when an assertion genuinely needs a mesh | + +A lab scenario takes tens of seconds and can fail for reasons that have nothing to do with the +artifact, and a shared-library change produces a cascade of dozens. One expensive +non-deterministic gate fails in both directions: a flaky run marks a good artifact unfit, a lucky +one marks a bad artifact fit, and **neither failure looks like itself.** + +**A run that failed environmentally is not a verdict.** A machine that would not boot says nothing +about the artifact, and recording it as *unfit* is the same untruth as recording a dispatch as a +deploy. + +## What a result means + +> **The declaration is updated, and here is which nodes have applied it.** + +A pipeline does not wait for every node — one may be legitimately switched off for a week, and a +delivery mechanism that blocks on a sleeping laptop is one nobody will use. + +``` +delivered declaration updated for 5 nodes +applied 3 of 5 +outstanding 2 — last seen 4 days ago, 20 minutes ago +``` + +**Outstanding is not failure**, and conflating them is how the old system produced a stall with no +error anywhere. + +## What must exist first + +1. **The module graph, with build edges.** No graph, no rebuild set and no ordering. +2. **A recorded input closure per artifact**, so currency is answerable without building. +3. **Something that notices a reconciler is not converging.** Below. + +## Open, and the first is the real risk + +- **A loop that will not converge is harder to debug than a job that failed.** A failed job stops + and names its step; a reconciler retries forever. Without something that notices *this has been + trying for an hour*, the failure is **silence** — the fault this removes, reintroduced in a new + place. +- **The run identity people use is lost.** *Did my change go out?* is answerable today by opening + a pipeline. Something must replace that or this is worse to live with, whatever its properties. +- **Does a fit artifact declare itself?** If it does, merging to main deploys to production — + which may be wanted and is far too large a property to acquire by omission. +- **Rebuild storms are mostly behaviourally empty.** Reproducible builds would stop a cascade at + the first module whose output did not move; without them one commit redeploys the fleet for no + change in behaviour. +- **Detection stays the fragile input for latency**, though no longer for correctness. diff --git a/02-DECISIONS/0004-managed-files-are-generated-never-edited.md b/02-DECISIONS/0011-managed-files-are-generated-never-edited.md similarity index 92% rename from 02-DECISIONS/0004-managed-files-are-generated-never-edited.md rename to 02-DECISIONS/0011-managed-files-are-generated-never-edited.md index 5fdb421..6bcb1cd 100644 --- a/02-DECISIONS/0004-managed-files-are-generated-never-edited.md +++ b/02-DECISIONS/0011-managed-files-are-generated-never-edited.md @@ -1,17 +1,18 @@ --- +topic: building it status: accepted date: 2026-04-03 deciders: jochen reconstructed: true --- -# 4. Managed files are generated onto nodes and never edited there +# 11. Managed files are generated onto nodes and never edited there > Reconstructed after the fact from the evidence cited below. ## Context -[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) put every binding in the mesh +[ADR 0006](0006-the-substrate-and-the-control-plane.md) put every binding in the mesh database. But the things that consume those bindings — environment files, service definitions, daemon configuration, firewall rules — are files on a node's disk, because that is what the software reading them requires. @@ -26,7 +27,7 @@ directions at once. 1. **Bidirectional sync** — a node's edits flow back to the database. Rejected, and removed. Two writers and no arbiter: whichever synced last wins, and neither is authority. 2. **Files are authoritative; the database is a cache of them.** Rejected — it inverts - ADR 0003 and returns to state that cannot be reconciled across nodes. + ADR 0013 and returns to state that cannot be reconciled across nodes. 3. **Strictly one-directional: the database is written, files are generated.** Chosen. ## Decision diff --git a/02-DECISIONS/0011-the-installer-owns-linking.md b/02-DECISIONS/0011-the-installer-owns-linking.md deleted file mode 100644 index 9c1dfbc..0000000 --- a/02-DECISIONS/0011-the-installer-owns-linking.md +++ /dev/null @@ -1,60 +0,0 @@ ---- -status: superseded -superseded-by: 02-DECISIONS/0018-the-mesh-creates-no-symlinks.md -date: 2026-07-10 -deciders: jochen -reconstructed: true ---- - -# 11. The installer owns linking; nothing else creates a symlink - -> Reconstructed after the fact from the evidence cited below. The incident that earned the rule -> predates the record, and its date is not established here. - -## Context - -A service's definition lives in the module catalogue; its runtime directory and persistent data -live outside it. The mesh connects the two by linking the definition into the runtime location -— deliberately, so that runtime state and source stay separate while the running service reads -a current definition. - -A link is also the easiest thing in the world to create by hand while fixing something, and a -container engine resolves a bind mount through it. A hand-made link pointed a volume somewhere -it should not have, and **production data was lost**. - -## Considered options - -1. **Copy instead of linking.** Rejected. A copy goes stale silently, which trades data loss - for a service running a definition nobody can find. -2. **Allow links, document the hazard.** Rejected. The hazard is not knowable at the moment of - the mistake — the link looks right and the resolution happens inside the container engine. -3. **One component owns linking; everyone else is forbidden.** Chosen. - -## Decision - -The installer creates and repairs every link the mesh needs. It reconciles them: a missing -link is created, a stale one is repointed, and a real file found where a link belongs is -adopted into the node's override location and replaced. - -**Nothing else creates a symlink** — not a hook, not a fix, not an agent, not a person -debugging. The prohibition is absolute because the judgement required to make a safe exception -is exactly the judgement that was not available at the moment it mattered. - -## Consequences - -- The class of failure is closed, at the cost of a rule that reads as arbitrary to anyone who - has not seen the incident. That is why it is recorded here rather than only asserted. -- Links become reconcilable state rather than incidental filesystem facts. -- The rule is stated for humans and agents and is enforced by convention, not mechanism. A - check does not exist. -- The rule as written governs the mechanism rather than removing it. A link made by the - installer resolves the same way as one made by hand, so the hazard is narrowed and not - closed. [ADR 0018](0018-the-mesh-creates-no-symlinks.md) proposes widening this to "nothing - links, the installer included"; until that is accepted, this record governs. - -## References - -- Recorded as a non-negotiable in the governed constitution page, §2: *"Symlinks to repos or - service directories have caused production data loss via Docker volume path resolution. The - installer handles all linking. Never create symlinks manually."* -- Knowledge base: `services` — the reconciliation behaviour, including adoption of real files. diff --git a/02-DECISIONS/0018-the-mesh-creates-no-symlinks.md b/02-DECISIONS/0012-the-mesh-creates-no-symlinks.md similarity index 84% rename from 02-DECISIONS/0018-the-mesh-creates-no-symlinks.md rename to 02-DECISIONS/0012-the-mesh-creates-no-symlinks.md index 762c1f9..6d513a6 100644 --- a/02-DECISIONS/0018-the-mesh-creates-no-symlinks.md +++ b/02-DECISIONS/0012-the-mesh-creates-no-symlinks.md @@ -1,16 +1,16 @@ --- +topic: building it status: accepted date: 2026-08-23 deciders: jochen reconstructed: false -extends: 0011-the-installer-owns-linking.md --- -# 18. The mesh creates no symlinks — a derived file is a copy +# 12. The mesh creates no symlinks — a derived file is a copy ## Context -[ADR 0011](0011-the-installer-owns-linking.md) responded to production data loss — a hand-made +[ADR 0012](0012-the-mesh-creates-no-symlinks.md) responded to production data loss — a hand-made link, resolved through a container engine's volume handling, pointing a mount somewhere it should not have — by centralising linking in the installer and forbidding it everywhere else. @@ -25,7 +25,7 @@ Two things have changed since, and together they remove the argument that kept i silently while the catalogue moves on, so a link was the cheap way to guarantee the running node reads a current definition. That argument assumes the node's copy is unmanaged. -**It is not.** [ADR 0004](0004-managed-files-are-generated-never-edited.md) established that +**It is not.** [ADR 0011](0011-managed-files-are-generated-never-edited.md) established that everything on a node's disk is derived from the mesh and regenerated when its inputs change, and the installer already **reconciles** links rather than assuming them — repointing stale ones, adopting real files it finds where a link belongs. Reconciling content is the same @@ -34,12 +34,12 @@ operation as reconciling a pointer, plus a comparison. So the mesh already has the machinery that makes a copy safe, and is using a link to solve a problem that machinery solves better. Worse, a link is conceptually the wrong shape: it makes the node's runtime state a *pointer into source*, which is the one thing -[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) and ADR 0004 exist to prevent. +[ADR 0006](0006-the-substrate-and-the-control-plane.md) and ADR 0011 exist to prevent. State is derived onto nodes; it does not reach back. ## Considered options -1. **Keep ADR 0011 as the final position** — centralised linking, forbidden elsewhere. +1. **Keep ADR 0019 as the final position** — centralised linking, forbidden elsewhere. Rejected as the status quo. It governs the mechanism rather than removing it, and the failure it was written for remains reachable by any code path the installer trusts. 2. **Keep links but harden them** — canonicalise before mounting, refuse a link that escapes @@ -54,12 +54,12 @@ State is derived onto nodes; it does not reach back. **The mesh creates no symlinks.** A file a node needs is placed on that node as a real file, derived from the mesh and reconciled by the installer like every other managed file -([ADR 0004](0004-managed-files-are-generated-never-edited.md)). +([ADR 0011](0011-managed-files-are-generated-never-edited.md)). -The prohibition in ADR 0011 stands and widens: it ceases to be "only the installer may link" +The prohibition in ADR 0019 stands and widens: it ceases to be "only the installer may link" and becomes "nothing links, the installer included". -When this is accepted, ADR 0011 becomes superseded rather than edited — its reasoning is why +When this is accepted, ADR 0019 becomes superseded rather than edited — its reasoning is why the rule exists at all, and the incident behind it is the reason anyone believes either record. ## Consequences @@ -73,7 +73,7 @@ the rule exists at all, and the incident behind it is the reason anyone believes cost, and it is the whole cost: today a link cannot be stale, and a copy can. The answer has to be detection — the installer comparing what is on disk against what the mesh says should be — and it must be loud, because a silently stale definition is exactly the failure shape - this mesh keeps producing ([ADR 0008](0008-a-failed-step-fails-the-job.md)). + this mesh keeps producing ([ADR 0010](0010-delivery.md)). - Reconciliation gets more expensive: comparing content rather than checking a pointer's target, on every module, on every node. - Disk usage rises, trivially, and is not a consideration. @@ -90,12 +90,12 @@ the rule exists at all, and the incident behind it is the reason anyone believes - **Migration order.** Converting a node's links is a change to how its services resolve their own definitions, which is not a change to make everywhere at once. -Until those are answered this record stays `proposed`, and ADR 0011 remains the governing rule. +Until those are answered this record stays `proposed`, and ADR 0019 remains the governing rule. ## References -- [ADR 0011](0011-the-installer-owns-linking.md) — the incident, and the rule this widens. -- [ADR 0004](0004-managed-files-are-generated-never-edited.md) — the machinery that makes a +- [ADR 0012](0012-the-mesh-creates-no-symlinks.md) — the incident, and the rule this widens. +- [ADR 0011](0011-managed-files-are-generated-never-edited.md) — the machinery that makes a copy safe. - [`03-DESIGN/00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md) — what the installer does today, including reconciliation and adoption. diff --git a/02-DECISIONS/0013-an-artifact-is-build-output.md b/02-DECISIONS/0013-an-artifact-is-build-output.md deleted file mode 100644 index c14f0cd..0000000 --- a/02-DECISIONS/0013-an-artifact-is-build-output.md +++ /dev/null @@ -1,63 +0,0 @@ ---- -status: accepted -date: 2026-08-04 -deciders: jochen -reconstructed: true ---- - -# 13. An artifact is build output, never a source tree - -> Reconstructed after the fact from the evidence cited below. - -## Context - -A module is built once and deployed to every node assigned to it. What travels between those -two events is the artifact. - -For a long time the artifact was a filtered copy of the module's source directory. Deploying it -therefore meant resolving and installing its dependencies **on the target node** — which -requires the target to reach a package registry, at deploy time, for every node, every deploy. -A node with no route to the registry could not deploy code that had already been built -successfully. - -## Considered options - -1. **Ship source, install dependencies on the target.** Rejected — it is what existed. Deploy - becomes a network operation with a failure mode per node, and the code that runs is - assembled independently on each one. -2. **Ship source plus its resolved dependency tree.** Rejected: large, slow, and it ships the - dependency resolution's platform assumptions along with it. -3. **Ship a self-contained build output; a failed bundle fails the build.** Chosen. - -## Decision - -The artifact is the module's **build output directory** — compiled and bundled, with its -dependency graph inlined. Deploy is extract-and-run and touches no network. - -A build that cannot produce a self-contained output **fails**. It does not fall back to -shipping a dependency tree, because a fallback that works is a fallback that is never fixed — -an application of [ADR 0008](0008-a-failed-step-fails-the-job.md). - -## Consequences - -- A node can deploy without reaching a registry. What was built is what runs, identically, on - every node. -- Deploys are faster and their failure modes are local. -- **Everything not in the build output does not ship.** This is the decision's whole cost, and - it was paid several times before it was understood: migrations that read the source layout, - provisioning scripts that read the source layout, selection files never packaged at all. Each - worked in development, where the source is present, and silently did nothing after deploy. -- Any file a module needs at runtime must be deliberately placed into the build output. The - rule "the artifact is `dist/`" has to be applied to every file kind, not just compiled code, - and that generalisation was the expensive part. -- Bundling has its own failure modes that a compiler will not catch — a bundler can exit - successfully and produce output that cannot load. - -## References - -- `build: bundle artifacts so a deploy is extract-and-run` (#673), 2026-08-04. -- The consequences, in order: `Provision migrations and seeds read the source layout, not the - artifact` (#699), `Local migrations read the source layout too` (#700), both 2026-08-07. -- Knowledge base: `pipeline/artifacts-are-build-output`, `pipeline/bundling`, - `troubleshooting/shell-migrations-never-packaged`, `troubleshooting/flavors-never-packaged`, - `troubleshooting/esbuild-silent-tla-breakage`. diff --git a/02-DECISIONS/0006-schema-changes-are-numbered-migrations.md b/02-DECISIONS/0013-schema-changes-are-numbered-migrations.md similarity index 96% rename from 02-DECISIONS/0006-schema-changes-are-numbered-migrations.md rename to 02-DECISIONS/0013-schema-changes-are-numbered-migrations.md index f175dde..cb7519e 100644 --- a/02-DECISIONS/0006-schema-changes-are-numbered-migrations.md +++ b/02-DECISIONS/0013-schema-changes-are-numbered-migrations.md @@ -1,11 +1,12 @@ --- +topic: building it status: accepted date: 2026-05-14 deciders: jochen reconstructed: true --- -# 6. Schema and state changes are numbered migrations, in the same language as the code +# 13. Schema and state changes are numbered migrations, in the same language as the code > Reconstructed after the fact from the evidence cited below. diff --git a/02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md b/02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md deleted file mode 100644 index 2443f60..0000000 --- a/02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md +++ /dev/null @@ -1,78 +0,0 @@ ---- -status: accepted -date: 2026-08-04 -deciders: jochen -reconstructed: true ---- - -# 14. Build, publish and deploy are three silos with different cardinality - -> Reconstructed after the fact from the evidence cited below. - -## Context - -Delivery had been treated as one pipeline that a module passes through. It is not: its stages -run a different number of times. - -- Compiling happens **once per module feature**, on the build node. -- Packaging and uploading happens **once per module feature**, on the build node. -- Installing, configuring, starting and verifying happens **once per module feature per node**. - -Conflating them is what made earlier versions slow and hard to reason about. Work that should -happen once was being repeated per node, and the fan-out point was implicit rather than a -boundary anything could observe. - -The split had been declared before it was real. Packaging still happened inside the build, -which meant the boundary existed in the documentation and not in the code. - -## Considered options - -1. **One pipeline, stages that know their own cardinality.** Rejected — it is what existed. - Cardinality is then a property of each stage's implementation, and nothing can reason about - the pipeline as a whole. -2. **Two silos: build-and-publish, then deploy.** Rejected. It leaves packaging inside build, - so build must know every module, every feature, and how each composes its artifact — - exactly the coupling the split exists to remove. A failed upload then retries by re-sending - a stale package instead of re-packaging. -3. **Three silos, with an explicit handover between each.** Chosen. - -## Decision - -Delivery is three silos, and the boundaries are real: - -| Silo | Runs | Where | -|---|---|---| -| **build** | once per module feature | the build node | -| **publish** | once per module feature | the build node | -| **deploy** | once per module feature **per node** | every assigned node | - -Commands and events are addressed **per feature**, not per module. - -Build compiles and hands over a **staged tree** — not a package. Publish applies the module's -packaging rules, packages that tree, and uploads it. Publishing to a package registry *is* -publishing, so a module whose artifact is a package publishes in the publish silo, not the -build one. - -Modules are resolved into dependency **levels**, and a level completes before the next begins, -so a module always builds against its dependencies' freshly published versions. - -## Consequences - -- Work that should happen once happens once. The fan-out point is explicit and observable. -- A failed upload retries by re-packaging, because packaging belongs to the stage that - uploads. -- The handover is a staged tree in a known location rather than the build's working directory, - which is reference-counted and cannot be assumed to still exist when a later stage runs. -- The build node is now the only node that has already passed through two silos when the - fan-out happens. Anything tracking a node's stage must account for **both** pre-fan-out - stages; code that knew only about the first parked the build node forever while every other - node deployed cleanly. -- A recovery mechanism that knows a subset of the stages it guards is worse than none — it - reports success over a stall it cannot see. - -## References - -- `publish owns packaging — the silos were not actually split` (#677), 2026-08-04. -- Knowledge base: `pipeline/three-silos` — including the note that the older architecture - documents claimed otherwise and were stale until 2026-08-06. -- The build-node stage-tracking failure was observed on pipeline #5557. diff --git a/02-DECISIONS/0007-no-npm-workspace.md b/02-DECISIONS/0014-no-npm-workspace.md similarity index 94% rename from 02-DECISIONS/0007-no-npm-workspace.md rename to 02-DECISIONS/0014-no-npm-workspace.md index d1425d8..1140d37 100644 --- a/02-DECISIONS/0007-no-npm-workspace.md +++ b/02-DECISIONS/0014-no-npm-workspace.md @@ -1,11 +1,12 @@ --- +topic: building it status: accepted date: 2026-06-04 deciders: jochen reconstructed: true --- -# 7. No workspace — each module is a standalone package consuming published dependencies +# 14. No workspace — each module is a standalone package consuming published dependencies > Reconstructed after the fact from the evidence cited below. @@ -47,7 +48,7 @@ its dependencies' freshly published versions. - Development and the pipeline resolve imports identically. The divergence is gone by construction rather than by discipline. - A module in its own repository is not a special case. It builds exactly as a module in the - monorepo does — which is what makes [ADR 0010](0010-applications-live-in-their-own-repository.md) + monorepo does — which is what makes [ADR 0015](0015-applications-live-in-their-own-repository.md) cheap. - A cross-package change costs a publish-and-consume round trip. This is the real price, paid on every shared-library change. diff --git a/02-DECISIONS/0010-applications-live-in-their-own-repository.md b/02-DECISIONS/0015-applications-live-in-their-own-repository.md similarity index 92% rename from 02-DECISIONS/0010-applications-live-in-their-own-repository.md rename to 02-DECISIONS/0015-applications-live-in-their-own-repository.md index 324958c..8ad1396 100644 --- a/02-DECISIONS/0010-applications-live-in-their-own-repository.md +++ b/02-DECISIONS/0015-applications-live-in-their-own-repository.md @@ -1,11 +1,12 @@ --- +topic: building it status: accepted date: 2026-07-10 deciders: jochen reconstructed: true --- -# 10. Applications live in their own repository; the monorepo is for the mesh +# 15. Applications live in their own repository; the monorepo is for the mesh > Reconstructed after the fact from the evidence cited below. @@ -46,7 +47,7 @@ reject it. - An application's cadence is its own. It is not reviewed as mesh code and does not queue behind mesh work. - The separation is safe **only because** the pipeline and provisioning are identical either - side of it — which [ADR 0007](0007-no-npm-workspace.md) is what makes true. Without + side of it — which [ADR 0014](0014-no-npm-workspace.md) is what makes true. Without standalone packages this decision would fork the build. - The monorepo stops being an inventory of the installation, which is a precondition for publishing anything about it. @@ -59,6 +60,6 @@ reject it. - The rule is stated in the governed constitution page authored 2026-07-10, §3, as a convention violation reviewers must reject. -- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) decision 4 extends this from *new* +- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) decision 4 extends this from *new* applications to the modules already in the monorepo. - Knowledge base: `troubleshooting/unregistered-module-source`. diff --git a/02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md b/02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md deleted file mode 100644 index 9288c29..0000000 --- a/02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md +++ /dev/null @@ -1,138 +0,0 @@ ---- -status: accepted -date: 2026-08-22 -deciders: jochen -reconstructed: false ---- - -# 16. A lab node is a virtual machine running the real install - -## Context - -[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite -rather than a convenience: *"everything that manifests between nodes is discoverable only in -production, which is where every fault of 2026-08-22 was found."* - -Two things were measured while establishing what exists -([`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md)): - -- **There is no local mesh.** The dev tooling starts providers through the host's own init - system and reads credentials from host paths (`modules/hal/developer/tools/dev-env.ts:146-178`). - It borrows the machine because there is nowhere else to put a mesh. -- **The one containerised node in the repository has been unable to build since 2026-06-04**, - when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it. - -So the question is not how to improve a local mesh. It is what a node *is* when it is not a -physical machine. Every subsequent question — how faithful is faithful enough, what may be -mocked, which failures remain reachable — follows from that one answer. - -The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and -hardware virtualisation present. - -## Considered Options - -1. **An application container.** Rejected. **A node's job is to run containers**, so modelling - a node as one inverts the thing being modelled: module service stacks then require nested - containers through a privileged daemon, or a shared socket that makes isolation between - nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start - and the least like a node. - -2. **A system container.** Rejected, after first being recommended. It is genuinely good — - real init, properly nested containers, roughly a second to boot, cheap snapshots — and it - is the only option that makes a twenty-node run affordable. It was rejected because **the - scale requirement that justified it was invented rather than required**: the stated goal is - to run the real mesh, which is four nodes, on one computer. And a system container still - forces the question a virtual machine dissolves — *how faithful must a node be?* — which - then has to be answered again for every capability under test. - -3. **`systemd-nspawn`.** Rejected. Already present, so nothing to install, but too primitive: - no storage pools, no snapshot management, no network management, no virtual machines. - Snapshots are what make the loop fast, so the saving is not worth what it costs. - -4. **A virtual machine.** **Adopted.** A bare Arch Linux machine that the real install script - turns into a node. - -## Decision - -**A node in the mesh development lab is a virtual machine.** It boots a stock Linux image, -runs the real install, and becomes a node. It is not a model of a node, so no question arises -about how good the model is. - -The environment is called **the lab**. - -Three things follow directly and are decided here: - -### The lab is driven by `incus` - -Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the -bridges between them, through one interface. It also manages system containers, so if a run -ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite. - -Declared in `modules/hal/developer/module.yml`, so it installs the way every other package -does. - -### The simulated public segment uses TEST-NET-3 - -`203.0.113.0/24`, reserved by RFC 5737, never routable. - -This is not cosmetic. WireGuard decides per pair whether to write an `Endpoint` by testing the -peer's underlay address against an RFC1918 regex -(`modules/wireguard/hooks/index.ts:225-240`). A simulated public segment addressed from -private space makes the hub test as unreachable, so no spoke writes an endpoint for it, -nothing can initiate, **and the mesh silently never forms** — appearing as a WireGuard fault -rather than an addressing mistake. - -The production LAN subnet and the entire overlay address plan are reproduced unchanged. - -### The lab issues its own certificates - -Public names are certified by an ACME server inside the lab; `.internal` names keep the mesh -CA. **The lab keeps production's two-authority split rather than collapsing it**, because a -single-authority lab would hide any fault living in that split. - -This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a -published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate -failure rather than a mystery. - -## Consequences - -**The fidelity question disappears, and with it a class of argument.** There is no "how real -is this node" to litigate per capability, because the node is real. What remains not-real is a -short, enumerable list: the model provider, the public internet, and the public certificate -authority. - -**The install becomes the thing under test.** A container-shaped lab would have had to skip -the bootstrap entirely. Here it runs, so it is exercised on every fresh lab. - -**Reproducing the network is mostly a data problem.** The bootstrap performs no network -configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks -from mesh-DB rows. The lab therefore exercises the same code production runs rather than a -reimplementation ([`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md)). - -**Scale runs get expensive, and this is the real cost.** Four virtual machines are -comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out -reaching most consumers rather than all, a cascade that stalls with many modules — stay hard -to reproduce. The mitigation is that the same tooling runs system containers, so a scale run -remains possible at lower fidelity if one is ever genuinely needed. - -**Boot is slower, and it does not matter.** Ten to twenty seconds against roughly one. A run -includes a full delivery — build, publish, install, migrate — measured in minutes, so boot -time is noise. - -**One change is required before the lab can issue certificates.** The reverse proxy sets no -`caServer`, so it defaults to the public authority's *production* endpoint -(`modules/traefik/docker-compose.yml:17-19`). It must become configurable, defaulting to -production so real nodes are unaffected. Worth noting on its own: aiming at production rather -than staging means every certificate experiment on a real node consumes issuance quota. - -## References - -- [`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md) — what exists, and - the four host couplings that only obstruct a container-shaped node -- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology - being reproduced and the endpoint constraint -- [`03-DESIGN/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) — what the lab - is for -- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments - recording what it cost to get right -- RFC 5737 — reserved documentation address blocks diff --git a/02-DECISIONS/0016-the-lab.md b/02-DECISIONS/0016-the-lab.md new file mode 100644 index 0000000..4b0b6b0 --- /dev/null +++ b/02-DECISIONS/0016-the-lab.md @@ -0,0 +1,84 @@ +--- +topic: building it +status: accepted +date: 2026-08-28 +deciders: jochen +reconstructed: false +--- + +# 16. The lab + +*Consolidated 2026-08-28 from five records. The lab is one design and was split across five +decisions taken over three days; the reasoning is kept, the fragmentation is not.* + +The environment a change is run against before it reaches real machines. + +## A node in the lab is a virtual machine + +It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of +a node**, so no question arises about how good the model is — which is the whole reason for +paying the cost of virtual machines rather than containers. + +The lab is driven by **incus**, and a scenario is raised from a declaration. + +## A router is scenery, and is therefore a container + +**Nothing under test runs on a router.** It is not a participant, holds no identity, has nothing +installed on it by the mesh, and no assertion is ever made about its internals. It exists so that +packets between machines behave the way they behave in the world. + +The fidelity argument that makes a node a virtual machine does not reach it: what a router *is* +does not matter, only what it *does to traffic*. So a router is a system container, and the lab +is cheaper for it. + +## A scenario declares the underlay, and only the underlay + +**What a hosting provider and a home router would have provided**, before any of our software +touched the machine: + +- which segments exist, and their address ranges +- which machine sits on which segment, at which address +- what NAT sits between them, and which ports are forwarded through it +- which machines are detached, and may be attached or detached during a run + +**A scenario declares nothing about the overlay** — no overlay addresses, no hub, no peering, no +names, no certificates. Those are the mesh's job, and a scenario that supplied them would be +testing itself. + +> A scenario provides what a hosting provider and a home router would provide, and nothing our +> software is responsible for. + +## A scenario is a closed address space + +Every segment materialises as its own isolated link belonging to one scenario instance. **Two +scenarios raised from the same declaration hold the same addresses and never meet**, because +nothing joins their links. The declaration therefore keeps its literal addresses and they mean +exactly what they say. + +**The consequence that constrains everything else: the lab never reaches into a scenario over +IP.** It talks to a machine through the virtualisation layer's own channel — the way one would +use a console rather than the network. That is what makes two identical scenarios able to run at +once, and it is why placing anything inside a machine is a hypervisor operation rather than a +network one. + +## Two scenario classes, and the first has no pipeline + +| | **bootstrap** | **full** | +|---|---|---| +| contains | machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules | +| verdict from | what the host reports about the state it reconciled | a delivery result ending in verification | +| exercises | tiers 0 and 1 | tiers 2 and above, and modules | + +**The bootstrap class comes first**, because it is what develops the node host, and because a +full scenario needs tiers that do not exist yet. A lab that could only raise the larger class +would be a lab nobody could use until everything else was built. + +## Consequences + +- **The lab tests the real code path**, not a reimplementation of it. The network a scenario + produces is generated by the same code production runs. +- **Isolation is what makes it usable in parallel**, and it costs the ability to reach in over + IP. Everything the lab puts inside a machine — a binary, an image, a file — goes through the + hypervisor. +- **A sealed scenario cannot fetch anything**, which is a real limit rather than an inconvenience: + it is why images have to be placed and why a container runtime has to be in the base image. diff --git a/02-DECISIONS/0034-a-test-defends-a-decision.md b/02-DECISIONS/0017-a-test-defends-a-decision.md similarity index 98% rename from 02-DECISIONS/0034-a-test-defends-a-decision.md rename to 02-DECISIONS/0017-a-test-defends-a-decision.md index 492b7a0..57ce326 100644 --- a/02-DECISIONS/0034-a-test-defends-a-decision.md +++ b/02-DECISIONS/0017-a-test-defends-a-decision.md @@ -1,11 +1,12 @@ --- +topic: checking it status: accepted date: 2026-08-24 deciders: jochen reconstructed: false --- -# 34. A test defends a decision +# 17. A test defends a decision ## Context diff --git a/02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md b/02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md deleted file mode 100644 index 0da332b..0000000 --- a/02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md +++ /dev/null @@ -1,97 +0,0 @@ ---- -status: proposed -date: 2026-08-23 -deciders: jochen -reconstructed: false -extends: 0015-mesh-brokers-nodes-host-agents-think.md ---- - -# 17. Modules outside the platform core are grouped by domain, not by single function - -## Context - -[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) recomposes the platform's own modules -into bounded contexts named after their aggregates, and sends the rest out of the monorepo on -the grounds that they run *on* the mesh rather than being *of* it. - -That leaves the larger half unaddressed. Around three quarters of the catalogue are modules -that are neither part of the mesh's domain nor standalone applications: a firewall, a VPN, an -SSH daemon and a resolver; a file manager, a media player and a system monitor; a set of -media-library services. Today each is its own module, because one module is the unit of *one -piece of software*, and no other grouping exists. - -The result is that the catalogue's shape records what was installed, not what anything is for. -Four modules that together constitute "how a node is reachable" have no relationship the mesh -can see: they cannot be assigned, versioned, reasoned about or replaced as one thing, and a -change to how the mesh handles connectivity has to be made four times. - -This is the same failure ADR 0015 names for the core — *boundaries drawn by deployment accident -rather than by domain* — appearing outside it. - -## Considered options - -1. **Leave them as they are.** Rejected. The core gets domain boundaries and everything else - keeps accident boundaries, so the catalogue becomes harder to read after the refactor than - before it. -2. **One module per piece of software, with a tag or category field.** Rejected. A label is not - a boundary: it does not change what can be assigned, versioned or replaced as a unit, and it - drifts from the thing it labels. -3. **Group them into domain modules, each owning the software that serves one purpose.** - Proposed here. -4. **Extend ADR 0015's contexts to cover everything.** Rejected. Those contexts are named for - the mesh's own aggregates; a media library is not an aggregate of the mesh, and forcing it - into that model repeats the metaphor-naming mistake ADR 0015 exists to correct. - -## Decision - -*Proposed — the principle is settled; the domain list is not. See "Open" below.* - -Modules that are not part of the platform core are grouped into **domain modules**. A domain -is named for the concern it serves, and owns the software that serves it. The unit stops being -one piece of software and becomes one purpose. - -This extends ADR 0015 rather than replacing it. The eight bounded contexts for the mesh's own -domain stand unchanged. This decision covers what ADR 0015 leaves outside them. - -Naming follows the same rule as the core: **name the domain for what it does, not for what it -is made of**. Connectivity, not a VPN implementation. - -## Consequences - -- A domain becomes assignable, versionable and replaceable as one thing. Changing how nodes - reach each other is a change to one module. -- The catalogue's shape starts describing purpose. A reader can tell what a mesh is *for* from - its module list. -- Swapping an implementation stops being a module replacement, with the data-volume and - provisioning consequences that carries, and becomes a change inside a domain. -- The count drops sharply, which is a symptom of the improvement rather than the point of it. -- **Grouping conceals.** A domain module hides which implementation is in use, and every - operational question — which port, which unit, which credential — gains an indirection. -- The migration is not free and has no obvious increments: a domain is only useful once - everything belonging to it has moved. -- Some modules genuinely serve one purpose and are already correctly sized. Grouping for its - own sake would be the same error in the other direction. - -## Open - -**The domain list is not settled and this record does not invent one.** What is decided is the -principle; what is not decided is the set. Candidate groupings are visible in the catalogue — -connectivity and reachability, node presentation and desktop, media libraries, observation and -metrics, storage and data services — but naming them here would be reconstructing a decision -that has not been taken. - -Settling the list is a research effort, not an act of this record. Until it concludes, this -ADR stays `proposed`. That effort is -[`01-RESEARCH/005-domain-grouping`](../01-RESEARCH/005-domain-grouping/00-overview.md), and its -first measurement already narrows this record's scope: co-change analysis supports grouping for -reachability, argues against it for the provisioned infrastructure providers, and finds no -signal either way for the fifty modules that never change alongside anything. - -## References - -- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — the core decomposition this - extends, and its rule about naming a context after its aggregate. -- [ADR 0010](0010-applications-live-in-their-own-repository.md) — standalone applications are - already out of scope here; they are not domains and do not group. -- [`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md) - — the catalogue's current shape, which is the evidence for the problem. diff --git a/02-DECISIONS/0035-a-picture-is-read-from-what-runs.md b/02-DECISIONS/0018-a-picture-is-read-from-what-runs.md similarity index 93% rename from 02-DECISIONS/0035-a-picture-is-read-from-what-runs.md rename to 02-DECISIONS/0018-a-picture-is-read-from-what-runs.md index 15c7676..04d9d32 100644 --- a/02-DECISIONS/0035-a-picture-is-read-from-what-runs.md +++ b/02-DECISIONS/0018-a-picture-is-read-from-what-runs.md @@ -1,18 +1,19 @@ --- +topic: checking it status: accepted date: 2026-08-24 deciders: jochen reconstructed: false -extends: 0034-a-test-defends-a-decision.md +extends: 0017-a-test-defends-a-decision.md --- -# 35. A picture of a system is read from the system, never from what asked for it +# 18. A picture of a system is read from the system, never from what asked for it ## Context A scenario declaration is a file. A raised scenario is a set of machines, links and rulesets. The two are supposed to correspond, and the entire value of the lab rests on noticing when -they do not — [ADR 0034](0034-a-test-defends-a-decision.md) says a claim nothing checks is a +they do not — [ADR 0017](0017-a-test-defends-a-decision.md) says a claim nothing checks is a claim that will quietly stop being true. Drawing a scenario makes that concrete, and forces a choice that looks cosmetic and is not. @@ -91,8 +92,8 @@ difference read off directly. ## References -- [ADR 0034](0034-a-test-defends-a-decision.md) — a claim nothing checks stops being true. -- [ADR 0031](0031-the-lab-provides-the-underlay.md) — why the lab must not supply what the +- [ADR 0017](0017-a-test-defends-a-decision.md) — a claim nothing checks stops being true. +- [ADR 0016](0016-the-lab.md) — why the lab must not supply what the mesh is responsible for; the same instinct, applied to facts rather than to configuration. - [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the fault in production form. diff --git a/02-DECISIONS/0019-how-this-repository-works.md b/02-DECISIONS/0019-how-this-repository-works.md new file mode 100644 index 0000000..56cf26c --- /dev/null +++ b/02-DECISIONS/0019-how-this-repository-works.md @@ -0,0 +1,132 @@ +--- +topic: how we work +status: accepted +date: 2026-08-28 +deciders: jochen +reconstructed: false +--- + +# 19. How this repository works + +*Consolidated 2026-08-28 from ten records that were one decision seen from ten angles. The +reasoning is kept; the fragmentation is not.* + +## The repository + +**`novox/hq` is Novox's headquarters, and it is public.** + +Company-scoped, not the mesh's. Today almost everything in it is about the mesh, because the +mesh is what Novox is building — a fact about the present rather than a definition. A second +product would live here too. + +**Public** means written for a reader who is not its author and has no access to the mesh it +describes. Nothing here may contain routable addresses, real domain names, node names, absolute +paths, usernames or credentials. The test: *would this paragraph still teach a stranger running +an entirely different mesh?* + +**Separate from the code** because the cadence differs — a decision changes when thinking +changes, not when code changes — and because a public repository cannot be a private one's +subdirectory. + +**The naming rule:** a repository belonging to a product carries that product's prefix. A +company-scoped one does not. So this is `hq` and the mesh's are `mesh-*`. + +| Repository | Tier | Holds | +|---|---|---| +| `novox/mesh-host` | 0 | the node host — the one binary installed by hand | +| `novox/mesh-substrate` | 1 | the pinned tier-1 services, as declarations | +| `novox/mesh-control` | 2 | the control plane and its contexts | +| `novox/mesh-surfaces` | 3 | tools, web, cli — thin, no logic | +| `novox/mesh-sdk` | — | the mesh's own domain ([ADR 0009](0009-modules-and-the-graph.md)) | +| `novox/mesh-lab` | — | the lab: scenario lifecycle, networking, placement | +| `novox/hq` | — | this one | + +**The product is `Novox Mesh`**, shortened to `mesh` in internal use — repository names, the +module namespace, environment variables, paths. **`Nox` is an identity of Novox**, an agent +participant within the mesh's own model, not a second system. + +## The folders, and why they are numbered + +**The numbering is the flow.** Research produces a decision; the decision authorises a design. +Following the numbers walks the process in the order it happens. + +| | | +|---|---| +| `01-RESEARCH` | an open question, while it is open | +| `02-DECISIONS` | what was decided, and why | +| `03-DESIGN` | what is being built | +| `04-ISSUES` | something wrong at the level of design or governance | + +**`03-DESIGN` has two layers and they are never mixed.** `00-as-is/` describes the mesh that +exists, written from the implementation and the operational record. `01-to-be/` describes the one +being built toward. Every document says which it is. A statement about the future does not belong +in an as-is document, and an as-is document is never edited to describe an intention. + +**`04-ISSUES` is for design-level faults** — a rule enforced by nothing, a stated invariant that +is false, a failure the design permits to be silent. Not an operational ticket queue. + +## What a decision record is, and is not + +**If a decision is worth recording, it is worth a record. If it is not worth a record, it is not +recorded.** + +That bar has been read too generously. A *finding* is not a decision. A bug is not a decision. +**A record is warranted when there is a genuine fork**: a direction reversed, an alternative +seriously considered and likely to be proposed again, or something contested that needs to stay +settled. Everything else belongs in the design document, where the reasoning is read. + +**There is no ledger** — no separate document summarising, ranking or tracking decisions. A +chronological view is generated from frontmatter, which is what a ledger was actually for. + +**A number identifies a record and never changes.** It is not a position, and it cannot be +both — a position moves when the set changes, and an identity that moves is not one. + +That is not a preference. Records are referenced from **outside** this repository: code +comments, commit messages, the knowledge base. Renumbering once cost 96 references across two +code repositories, and nothing in either would have failed to compile — the comments would +simply have pointed at the wrong reasoning, which is worse than a broken link because nothing +reports it. + +**So the reading order lives in a generated index**, from each record's `topic:` — what the mesh +is, then its tiers from the bottom up, then what runs on them and how it gets there, then how it +is built, how it is checked, and how we work. + +**And the index is written, not only generated on demand.** A reader looking at the folder on a +forge sees the folder, not a command. The objection to a written index is that it drifts, and +that is answered by **checking** it rather than by refusing to write one — which is §5's own +rule: a rule states how it is checked. A record with no topic, or a topic nobody defined, fails +the same check, because the quiet failure is a record that vanishes from the order rather than +appearing in the wrong place. + +**The design layer is what you read.** These records explain *why* a thing is as it is. They are +not a description of the system, and needing to read them to understand it would mean the design +documents had failed. + +## Status, and views over it + +**Every document carries its state in YAML frontmatter** — research overviews, design documents, +decision records, issue reports. + +**There are no central status files.** Every cross-cutting view — a status matrix, a decision +index, an open-issue list — is generated from frontmatter when asked for, never written to disk. +Two places holding one fact drift, and the written one wins by being closer to hand. + +**Prose does not restate status.** One place, and two is one too many. + +## Workflows are playbooks + +Every workflow is a playbook in [`00-META/process/`](../00-META/process/): trigger, who runs it, +steps, outputs. People and agents follow the same ones, and **agents do not act outside them**. + +Each is wrapped by a thin skill that defers to the playbook as authoritative and adds only the +mechanical scaffolding — so the process has one definition rather than a document and an +implementation that disagree. + +## Consequences + +- **A reader has one place per thing.** The design layer describes the system; these records say + why; the playbooks say how work is done. +- **Records will accumulate more slowly**, because the bar is a fork rather than a finding. This + record is itself the correction: ten records became one because they were one decision. +- **The public rule constrains everything written here**, permanently and at every commit. It is + the reason research describes real observations without identifying the mesh it observed. diff --git a/02-DECISIONS/0019-hq-is-its-own-repository.md b/02-DECISIONS/0019-hq-is-its-own-repository.md deleted file mode 100644 index f712df4..0000000 --- a/02-DECISIONS/0019-hq-is-its-own-repository.md +++ /dev/null @@ -1,69 +0,0 @@ ---- -status: accepted -date: 2026-08-22 -deciders: jochen -reconstructed: false -supersedes: none ---- - -# 19. HQ is its own repository, and it is public - -## Context - -The mesh's reasoning — mission, research, design, decisions — began inside the code -repository, under a folder there. The objection to moving it out was specific and good: the -mesh already has an operational memory and a structured archive, and adding a third store -repeats the mistake that consolidation was meant to fix. - -## Considered options - -1. **Keep it in the code repository.** Rejected, but the objection it rests on is correct and - is answered rather than dismissed — see Consequences. -2. **Put it in the structured archive**, alongside the governed documents. Rejected: the - archive is not reviewable as a diff, and a design argument is exactly the thing that needs - line-by-line review and a branch. -3. **Its own repository.** Chosen. - -## Decision - -HQ is its own repository, and it is **public** — written for a reader who is not its author -and has no access to the mesh it describes. - -Three reasons it is separate: - -- **The cadence differs.** A decision changes when thinking changes, not when code changes. - Tying documents to a code branch merges them on the code's schedule. -- **The reviewers differ.** A design argument is not reviewed the way an implementation is, - and should not queue behind a build. -- **The scope is wider than one repository.** - [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) sends most modules out of the - monorepo; documentation governing several repositories cannot live inside one of them. - -Being public is not incidental. It is enforceable only because -[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) made the code repository -node-agnostic: there is no per-node content to leak. Nothing here may carry routable -addresses, real domain names, hosting providers, node names, absolute paths, usernames, -credentials, or operational detail useful only to an attacker. - -The test is whether a paragraph would still teach a stranger running an entirely different -mesh. - -## Consequences - -- A document and the code it describes can no longer land in one commit. Keeping them honest - is a discipline rather than a mechanism — which is why decisions are recorded as they are - taken, and why a document stating a rule must say how the rule is checked. -- Research must state evidence without identifying the mesh it observed. The shape of a - finding survives anonymisation; the instance does not travel. -- **The objection is answered by indexing, not by location** — the claim being that these - documents remain searchable beside everything else, one source with many surfaces. - **That indexing does not exist.** Checked 2026-08-23, it returns nothing. Until it does, the - objection stands unanswered and this repository is the third knowledge store it was argued - not to be. Recorded as - [`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md). - -## References - -- Supersedes the earlier position that documentation lives inside the code repository under a - folder there. That position was never recorded separately and has no record of its own. -- [`README.md`](../README.md) — the public-repository rule in full. diff --git a/02-DECISIONS/0020-design-is-written-in-two-layers.md b/02-DECISIONS/0020-design-is-written-in-two-layers.md deleted file mode 100644 index cf671bc..0000000 --- a/02-DECISIONS/0020-design-is-written-in-two-layers.md +++ /dev/null @@ -1,67 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 20. Design is written in two layers: what is, and what is intended - -## Context - -HQ held only the intended mesh. Every reader had to already know the running system in order -to understand what the decisions were about, and a statement about current behaviour had -nowhere to live except inside a document describing an intention. - -The consequence was invisible until looked for: an as-is claim inside a to-be document is -indistinguishable from the intention around it, so the document silently stops being true as -the system moves — and nobody can tell which half went stale. - -The mesh also has a large body of shipped behaviour that nobody would choose again. It is not -design in the sense of "what we decided"; it is design in the sense of "what is there", and it -is exactly the part a person changing the system most needs. - -## Considered options - -1. **One layer, describing the target.** Rejected — the status quo. The running system goes - undocumented and the target document accumulates unmarked claims about it. -2. **One layer, describing what runs, with intentions only in decision records.** Rejected: - a decision record is an argument, not a specification, and a multi-part intention has - nowhere coherent to live. -3. **Two layers, declared per document, never mixed.** Chosen. - -## Decision - -`03-DESIGN` holds two layers, and every document declares which it is: - -| Layer | Describes | Written from | -|---|---|---| -| `00-as-is/` | The mesh that exists | The implementation and the operational record | -| `01-to-be/` | The mesh being built toward | Decision records | - -An as-is document **records what is, not what should be** — including behaviour nobody would -choose again. A layer that keeps only the good decisions is a brochure. - -When a to-be design ships it **does not move**. Its as-is counterpart is written or updated, -the to-be document's status becomes `implemented`, and both stand: one describing what runs, -the other recording what was intended. Deleting the intention loses the reasoning. - -Where implementation and intention disagree, the as-is document records the implementation and -says they disagree. - -## Consequences - -- A reader can tell, from the folder and from one frontmatter field, whether they are reading - a description or a plan. That distinction was previously unavailable at any price. -- Correcting an as-is document requires evidence from the implementation, not agreement — and - needs no decision record, because nothing was decided. -- Two documents must be kept current per subsystem instead of one. This is the cost, and it is - paid on every ship. -- Something that shipped differently from its design becomes a visible divergence rather than - a silently wrong document, and may deserve an issue. - -## References - -- [`03-DESIGN/README.md`](../03-DESIGN/README.md) — the layer contract and frontmatter schema. -- [`03-DESIGN/00-as-is/`](../03-DESIGN/00-as-is/) — the first eleven as-is documents, written - 2026-08-23 from the monorepo and the operational memory. diff --git a/02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md b/02-DECISIONS/0020-the-mesh-is-governed-by-a-constitution.md similarity index 96% rename from 02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md rename to 02-DECISIONS/0020-the-mesh-is-governed-by-a-constitution.md index 0088030..c43cbdd 100644 --- a/02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md +++ b/02-DECISIONS/0020-the-mesh-is-governed-by-a-constitution.md @@ -1,11 +1,12 @@ --- +topic: how we work status: accepted date: 2026-07-10 deciders: jochen reconstructed: true --- -# 9. The mesh is governed by a constitution, injected where work is decided +# 20. The mesh is governed by a constitution, injected where work is decided > Reconstructed after the fact from the evidence cited below. diff --git a/02-DECISIONS/0025-hq-is-the-source-of-the-constitution.md b/02-DECISIONS/0021-hq-is-the-source-of-the-constitution.md similarity index 90% rename from 02-DECISIONS/0025-hq-is-the-source-of-the-constitution.md rename to 02-DECISIONS/0021-hq-is-the-source-of-the-constitution.md index 4028902..9317c7a 100644 --- a/02-DECISIONS/0025-hq-is-the-source-of-the-constitution.md +++ b/02-DECISIONS/0021-hq-is-the-source-of-the-constitution.md @@ -1,15 +1,16 @@ --- +topic: how we work status: accepted date: 2026-08-23 deciders: jochen reconstructed: false --- -# 25. HQ is the source of the mesh constitution +# 21. HQ is the source of the mesh constitution ## Context -[ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) established a canonical rule set, +[ADR 0020](0020-the-mesh-is-governed-by-a-constitution.md) established a canonical rule set, injected into every eligible design session and checked before output is accepted. It lives in the knowledge base, where the orchestrator reads it. @@ -39,7 +40,7 @@ directly. Publishing is a playbook step, not a manual act, and it ends with **reading the page back and verifying the change is present**. A publish that reported success and did nothing is exactly the failure class this mesh keeps producing -([ADR 0008](0008-a-failed-step-fails-the-job.md)). +([ADR 0010](0010-delivery.md)). Section numbering is stable, because the orchestrator and the review fragments cite sections by number. @@ -47,7 +48,7 @@ number. ## Consequences - One source, many surfaces — the same argument HQ's separation already rests on - ([ADR 0019](0019-hq-is-its-own-repository.md)), applied to the rules themselves. + ([ADR 0019](0019-how-this-repository-works.md)), applied to the rules themselves. - Each rule keeps the incident that earned it, in a place that is reviewed as a diff. - An edit to the derived page survives until the next sync and then vanishes. The playbook says so, and nothing mechanically prevents it. @@ -63,5 +64,5 @@ number. - [`00-META/process/05-constitution-sync.md`](../00-META/process/05-constitution-sync.md) — the sync, including the read-back. -- [ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) — the governed page and why it +- [ADR 0020](0020-the-mesh-is-governed-by-a-constitution.md) — the governed page and why it exists. diff --git a/02-DECISIONS/0021-work-moves-through-playbooks.md b/02-DECISIONS/0021-work-moves-through-playbooks.md deleted file mode 100644 index 9618c19..0000000 --- a/02-DECISIONS/0021-work-moves-through-playbooks.md +++ /dev/null @@ -1,57 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 21. Every workflow is a playbook, and agents operate through them - -## Context - -HQ stated a knowledge flow — research becomes design — and nowhere stated how anything moves -along it. What graduation required, who wrote the decision, what closed an effort, what -happened when something shipped: all of it was convention held in one person's head. - -A large share of the work here is done by agents. An unwritten convention is not available to -an agent at all, so each one either invents a procedure or asks. Both produce a repository -whose shape depends on who last touched it. - -## Considered options - -1. **Convention, learned by reading existing documents.** Rejected — the status quo. It - transmits shape but not rules, and it transmits the mistakes along with the pattern. -2. **One long contributing document.** Rejected: it is read once, and the step someone needs - is never the step they are reading. -3. **A playbook per workflow, each with trigger, steps and outputs, wrapped by a thin skill.** - Chosen. - -## Decision - -Every workflow is a playbook in [`00-META/process/`](../00-META/process/): trigger, who runs -it, steps, outputs. Engineers and agents follow the same playbooks, and **agents must not act -outside them**. - -Each playbook is wrapped by a thin skill that defers to it as authoritative and adds only -mechanical scaffolding — next free number, frontmatter block, where the file goes. The -playbook holds the reasoning; the skill holds the steps. When they disagree, the playbook -wins. - -## Consequences - -- An agent arriving with no context can act correctly, because the procedure is retrievable - rather than remembered. -- The playbooks are themselves reviewable. A bad rule can be found and changed, which is not - true of a convention. -- Duplication between playbook and skill is real, and is managed by making the skill thin and - naming the playbook as authoritative in the skill's first lines. Nothing prevents them - drifting; the constraint is that only one carries reasoning. -- A workflow with no playbook is a workflow agents will get wrong. Adding one is part of - adding the workflow. - -## References - -- [`00-META/process/00-overview.md`](../00-META/process/00-overview.md) — the five playbooks - and the flow they implement. -- Modelled on the process layer in the sibling HQ repository for the PAPA platform, which - arrived at the same shape and the same thin-skill split. diff --git a/02-DECISIONS/0022-status-lives-in-frontmatter.md b/02-DECISIONS/0022-status-lives-in-frontmatter.md deleted file mode 100644 index cb99f15..0000000 --- a/02-DECISIONS/0022-status-lives-in-frontmatter.md +++ /dev/null @@ -1,54 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 22. Status lives in frontmatter; cross-cutting views are generated - -## Context - -Status was carried in prose — a bold line near the top of a document saying what state it was -in — and indexes were maintained by hand. The decision-record index had already drifted from -the folder it described **after a single addition**, which is about as short a demonstration -as the failure mode offers. - -A hand-maintained index is a copy of something the filesystem already knows. It is correct -only for as long as everyone remembers it exists, and its being wrong is silent. - -## Considered options - -1. **Prose status plus hand-maintained indexes.** Rejected — the status quo, already - demonstrably broken. -2. **A central status file.** Rejected. It centralises the drift rather than removing it: the - file and the documents disagree, and the file is the one people read. -3. **Machine-readable frontmatter per document; every cross-cutting view generated on - demand.** Chosen. - -## Decision - -Every document carries its state in YAML frontmatter — research overviews, design documents, -decision records, issue reports — with a schema stated in the section README. - -**There are no central status files.** Every cross-cutting view — a status matrix, the -decision-record index, the open-issue list — is generated from frontmatter when asked for, and -never written to disk. - -Prose does not restate status. One place, and two is one too many. - -## Consequences - -- A view cannot drift from what it describes, because it does not persist. -- Status becomes queryable. Inconsistencies — a closed effort with nothing in `became:`, an - `implemented` design with no owning repository — are findable mechanically, and the - generator reports them as flags rather than silently rendering around them. -- Frontmatter must be valid and paths in it must resolve, which is now something to check. -- A reader browsing the repository on a forge sees no index. That is the trade: the index is - correct and absent rather than present and wrong. - -## References - -- [`.claude/skills/hq-status/SKILL.md`](../.claude/skills/hq-status/SKILL.md) — the - generator, including the inconsistencies it flags. -- [`02-DECISIONS/README.md`](README.md) — the hand-written index that drifted, and its removal. diff --git a/02-DECISIONS/0040-the-constitution-absorbs-what-is-enforced.md b/02-DECISIONS/0022-the-constitution-absorbs-what-is-enforced.md similarity index 92% rename from 02-DECISIONS/0040-the-constitution-absorbs-what-is-enforced.md rename to 02-DECISIONS/0022-the-constitution-absorbs-what-is-enforced.md index a0e9508..7b7a608 100644 --- a/02-DECISIONS/0040-the-constitution-absorbs-what-is-enforced.md +++ b/02-DECISIONS/0022-the-constitution-absorbs-what-is-enforced.md @@ -1,15 +1,16 @@ --- +topic: how we work status: accepted date: 2026-08-26 deciders: jochen reconstructed: false --- -# 40. The constitution absorbs what is already enforced +# 22. The constitution absorbs what is already enforced ## Context -[ADR 0025](0025-hq-is-the-source-of-the-constitution.md) makes this repository the source +[ADR 0021](0021-hq-is-the-source-of-the-constitution.md) makes this repository the source and the knowledge base a derived copy, and playbook [05](../00-META/process/05-constitution-sync.md) publishes the copy whenever a rule changes. @@ -97,8 +98,8 @@ trusting the second success message either. ## References -- [ADR 0025](0025-hq-is-the-source-of-the-constitution.md) — source and copy. -- [ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) — why the copy is injected at all. +- [ADR 0021](0021-hq-is-the-source-of-the-constitution.md) — source and copy. +- [ADR 0020](0020-the-mesh-is-governed-by-a-constitution.md) — why the copy is injected at all. - [Playbook 05](../00-META/process/05-constitution-sync.md) — the sync this record interrupts. -- [ADR 0034](0034-a-test-defends-a-decision.md), [ADR 0035](0035-a-picture-is-read-from-what-runs.md), - [ADR 0018](0018-the-mesh-creates-no-symlinks.md) — the three rules whose sync surfaced this. +- [ADR 0017](0017-a-test-defends-a-decision.md), [ADR 0018](0018-a-picture-is-read-from-what-runs.md), + [ADR 0012](0012-the-mesh-creates-no-symlinks.md) — the three rules whose sync surfaced this. diff --git a/02-DECISIONS/0042-approval-is-the-checkpoint.md b/02-DECISIONS/0023-approval-is-the-checkpoint.md similarity index 95% rename from 02-DECISIONS/0042-approval-is-the-checkpoint.md rename to 02-DECISIONS/0023-approval-is-the-checkpoint.md index a4e25fa..2ec3fad 100644 --- a/02-DECISIONS/0042-approval-is-the-checkpoint.md +++ b/02-DECISIONS/0023-approval-is-the-checkpoint.md @@ -1,11 +1,12 @@ --- +topic: how we work status: accepted date: 2026-08-26 deciders: jochen reconstructed: false --- -# 42. The approval is the checkpoint, not the second pair of hands +# 23. The approval is the checkpoint, not the second pair of hands ## Context @@ -73,5 +74,5 @@ the previous sync reported success and changed nothing. ## References - [`how-we-build.md`](../00-META/how-we-build.md) §2 — the rule, now carrying this. -- [ADR 0008](0008-a-failed-step-fails-the-job.md) — the standard a checkpoint is held to: a +- [ADR 0010](0010-delivery.md) — the standard a checkpoint is held to: a step that reports success without doing anything is the fault, not the shortcut. diff --git a/02-DECISIONS/0023-issues-have-a-front-door.md b/02-DECISIONS/0023-issues-have-a-front-door.md deleted file mode 100644 index e844595..0000000 --- a/02-DECISIONS/0023-issues-have-a-front-door.md +++ /dev/null @@ -1,60 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 23. Issues have a front door, separate from the operational memory - -## Context - -Findings that were nobody's task accumulated in a table inside the decision ledger — a package -install reporting success while installing nothing, a manifest key read by no code, an -end-to-end harness dead for months. They were measured, true, and unowned: a table row cannot -be assigned, diagnosed or closed. - -The mesh already has an operational memory holding roughly a hundred and thirty entries, -indexed on symptoms. The obvious move — put these there — is wrong, and the reason is the -distinction worth recording. - -## Considered options - -1. **Leave them in the ledger.** Rejected: a ledger records decisions taken, and these are - the opposite — questions nobody has answered. -2. **Put them in the operational memory.** Rejected. That store answers *how do I fix this - occurrence*; these are *why does the design allow this at all*. Filing them there makes - them findable by symptom and unfindable as open questions, and nothing there has a state - that can be closed. -3. **A numbered issue folder in HQ, deliberately narrow.** Chosen. - -## Decision - -`04-ISSUES` is the front door for something wrong at the level of **design or governance**: -a rule enforced by nothing, a stated invariant that is false, a failure the design permits to -be silent, or a symptom whose owner cannot be found without the whole mesh in view. - -One numbered folder per issue: the report with the symptom as observed and the evidence, and -a diagnosis document carrying the trail, dated, including what was ruled out. - -**This is not a second copy of the operational memory.** An issue here is a question HQ must -*answer*; an entry there is an incident someone must *clear*. An issue whose answer is a -general lesson belongs in both — and the operational memory is searched first, because if the -answer is already there this was never an issue. - -## Consequences - -- A finding gets a number, a state and an owner, and closing it is a visible act. -- The symptom-to-component trail accumulates in a place where the whole mesh is visible, which - is where cross-component diagnosis has to happen. -- The boundary needs judgement on every report, and will sometimes be got wrong. Filing too - narrowly loses a finding; filing too widely rebuilds the operational memory here, which is - the outcome HQ's separation was argued against - ([ADR 0019](0019-hq-is-its-own-repository.md)). -- Six issues opened on creation, all previously unowned observations. - -## References - -- [`04-ISSUES/README.md`](../04-ISSUES/README.md) — the boundary table and the frontmatter - schema. -- [`00-META/process/03-issues.md`](../00-META/process/03-issues.md) — the playbook. diff --git a/02-DECISIONS/0024-model-access-is-a-provision.md b/02-DECISIONS/0024-model-access-is-a-provision.md new file mode 100644 index 0000000..c40bbb7 --- /dev/null +++ b/02-DECISIONS/0024-model-access-is-a-provision.md @@ -0,0 +1,95 @@ +--- +topic: what runs on it +status: accepted +date: 2026-08-30 +deciders: jochen +reconstructed: false +--- + +# 24. Model access is a provision, and a licence is a thing with a name + +## Context + +Everything in this mesh that thinks needs a model, and there is more than one way to reach one: + +| | | +|---|---| +| **hosted services** | several vendors, each with its own account, quota and key | +| **models the mesh runs itself** | open-weight models on a node with the hardware for them | + +And the choice is **per consumer, deliberately**: a workstation's own session on one account, a +laptop on the organisation's, two hired workers on the mesh's local model because their work does +not justify paid tokens. Those are three different answers to one requirement, held at once, in +one mesh. + +**The existing system has the hard half of this already** — automatic licence refresh and +switching between accounts when one is exhausted — and it works. It is not being replaced because +it was wrong; it is being rebuilt because it lives in a place that cannot express the rest. + +## Decision + +**Model access is a provision.** A module that needs to think declares `requires: model-access`; +anthropic, openai, grok and a locally-run model are four modules that provide it. That is +[ADR 0009](0009-modules-and-the-graph.md)'s mechanism unchanged, and it buys the things that +mechanism already buys: several implementations of one job, a refusal when more than one could +answer, and choosing by assigning the one you want. + +**A locally-run model needs nothing new at all.** It is a mesh-scoped provision on the node with +the hardware — the same shape as a database, including the credential. + +### A licence is a named thing, and the name is the operator's + +Not an anonymous credential hanging off a provider. *The personal account*, *the organisation's +account* — those are names a person uses, and the mesh has to use them too, because the whole +point is saying **which one** a given consumer uses. + +**Many to many.** One provider has several licences; one licence serves several consumers. So it +is **not a claim** — claims are for things only one holder may have, and two machines sharing an +account is the ordinary case rather than a collision. + +### Four things this needs that the mesh does not have + +Written as gaps rather than as design, because each is a real piece of work and pretending +otherwise is how a plan becomes a surprise. + +**1. A provider that is not on a node.** A mesh-scoped provision today is answered by *the machine +running it*, and the reachability rule refuses two ends that share no private network. A hosted +service is on nobody's machine and is reached over the public internet. That is a third scope — +answered by a record rather than by a node — and the reachability rule must not apply to it. + +**2. A secret the mesh is given rather than one it mints.** Every credential the mesh handles +today it generated itself, sealed to both ends, and discarded. An API key arrives from a person. +The missing verb is *accept*: take a value, seal it to each holder, and **discard the plaintext** +— because a mesh that keeps operator-supplied keys readably is the arrangement this project +[measured and rejected](0009-modules-and-the-graph.md). + +**3. A consumer that is not a machine.** *This worker uses that licence* is a binding to an agent, +not to a node. [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) already says an agent +holds credentials and that delivery follows its node bindings and modality — so what is delivered +still lands on a machine, and what is **chosen** is chosen per agent. The provisions model has no +consumer identity other than a node. + +**4. Switching is a reaction, not a declaration.** Everything here is desired state, reconciled by +comparison. A licence that hits its limit and must be swapped is a response to something observed, +and it cannot be expressed as a declaration without the declaration meaning *whatever is working +right now* — which is not a thing anybody declared. **It belongs with observability, changing a +binding**, and the binding is then declared as usual. Saying this plainly is what stops the +declaration language growing a conditional. + +## Consequences + +- **The refusing rule applies here and will be felt.** A mesh holding three ways to reach a model + refuses every consumer that has not said which — which is correct and is a great deal of + saying-which the first time. The remedy is one assignment per consumer, and the message names + the candidates. +- **A licence outliving its holder is a live credential nobody is watching.** The same rule the + provisioner follows applies: what the mesh granted and no longer grants is withdrawn. +- **Nothing here makes a node authenticate to a model provider.** ADR 0001 holds: an agent does. + What changes is that the mesh can now say *which agent, which licence, which node it lands on*, + which is the fact ADR 0001 records as missing. + +## References + +- [ADR 0009](0009-modules-and-the-graph.md) — provisions, scope, choosing, and sealed credentials +- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — agents hold credentials, not nodes; + `hal/ai` as the context owning provider grants and rotation diff --git a/02-DECISIONS/0024-the-numbering-is-the-flow.md b/02-DECISIONS/0024-the-numbering-is-the-flow.md deleted file mode 100644 index b1626d0..0000000 --- a/02-DECISIONS/0024-the-numbering-is-the-flow.md +++ /dev/null @@ -1,76 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 24. The folder numbering is the flow, and decision records run oldest first - -## Context - -Two orderings were wrong in ways that only show up when someone new reads the repository. - -**The folders.** Decisions lived in an unnumbered folder that sorted after the numbered ones, -so the repository's most load-bearing content read as an annex. - -The sibling HQ repository for the PAPA platform had already solved this and solved it -crookedly: its design folder existed from its initial commit, and when its decision folder was -finally promoted it took the **next free number** rather than its place in the sequence. That -repository now reads `01 research → 03 decision → 02 design`. A decision precedes the design -it authorises and is numbered after it. By the time this was visible, the design folder was too -settled to renumber. - -**The records.** Fourteen decisions had been taken in implementation and never written down — -the broker, the module abstraction, the mesh database, the artifact, the three silos and the -rest. Meanwhile two records existed, holding numbers 0001 and 0002, for decisions taken last. - -## Considered options - -1. **Match the sibling repository exactly**, inheriting its ordering. Rejected: structural - parity is worth something, but not the cost of copying a scar the other repository would - not choose again. -2. **Leave the folder unnumbered.** Rejected — the annex problem, and it leaves an unexplained - gap for anyone arriving from the sibling repository. -3. **Number by position in the flow, and renumber the records chronologically.** Chosen, on - the grounds that this repository was four commits old and nothing outside it cited a - number. That is the only window in which either renumbering is free. - -## Decision - -**The numbering is the flow.** Research produces a decision; the decision authorises a design. -So `01-RESEARCH`, `02-DECISIONS`, `03-DESIGN`, `04-ISSUES`. Following the folder numbers walks -the process in the order it happens. - -**Decision records are a chronological ledger.** They run oldest first. The fourteen decisions -already taken in implementation were back-filled as records 0001–0014, each dated from the -history, each carrying `reconstructed: true` and saying so in its first lines, and each citing -the commit, pull request or knowledge-base entry it was recovered from. The two existing -records moved to 0015 and 0016. - -A reconstructed record is not a transcript. Where the deliberation is not recoverable it states -what the alternatives were and why the chosen one won on the evidence available — not a -discussion that did not happen. Where a date is not establishable it says so. - -The foundational folder is `00-META`, matching the sibling repository. - -## Consequences - -- The repository reads in process order, and the gap at `03` that a reader coming from the - sibling repository would notice is explained by this record. -- The design documents can cite reasoning instead of asserting rules, because the reasoning now - exists. -- Structural divergence from the sibling repository, deliberately, in exactly one place. It is - recorded here so that the difference reads as a choice rather than an accident. -- **Record numbers are now stable and renumbering is over.** This decision spends the one - window that existed; a future record takes the next free number regardless of its date. -- Reconstructed records carry a standing risk: they are the most confident-sounding documents - in the repository and the least directly witnessed. The `reconstructed` flag exists so that - is never invisible. - -## References - -- The sibling repository's restructure of 2026-07-13 moved its decision folder in a single - commit of twelve renames with no content change, alongside the same status-into-frontmatter - and playbook changes made here. -- [`02-DECISIONS/README.md`](README.md) — the format, and the note on reconstructed records. diff --git a/02-DECISIONS/0025-the-design-record-is-read-not-copied.md b/02-DECISIONS/0025-the-design-record-is-read-not-copied.md new file mode 100644 index 0000000..7d4e5b3 --- /dev/null +++ b/02-DECISIONS/0025-the-design-record-is-read-not-copied.md @@ -0,0 +1,105 @@ +--- +topic: how we work +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0019-how-this-repository-works.md +--- + +# 25. The design record is read where it is written, never copied to be found + +## Context + +**These documents cannot be found by searching the mesh's memory, and never could.** Checked on +2026-08-23 and again on 2026-08-31, against both the symptom-indexed store and the structured +archive, using a decision record's full title and a distinctive phrase from a design document: no +result, no partial match, no stale copy. + +That matters because of what was promised. The objection to giving this material its own +repository was that the mesh already has a knowledge store, and a second one repeats the mistake +that store was created to fix. **The answer offered was indexing rather than location** — that +these documents would be returned beside everything else in a search, so where they were authored +became a separate question. The indexing was never built. + +**The claim has since stopped being load-bearing**, which is why this is a decision rather than an +incident. [`README.md`](../README.md) names the gap in the place the claim used to sit, and +[ADR 0019](0019-how-this-repository-works.md)'s reasoning rests on cadence, reviewers and scope — +none of which depend on being searchable from elsewhere. What remained was an unbuilt capability +and an open question, recorded as +[`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md). + +**A signpost was added on 2026-08-31 and measured.** One entry in the mesh's memory naming what +lives here and when to come looking. A search for *design records, decisions, repository* returns +it; a search phrased the way somebody actually asks — *why is the mesh built this way* — returns +nothing, because the store matches terms and not meaning. **Reachable is not the same as +surfacing**, and the measurement is what established which one a signpost buys. + +## Considered Options + +1. **A one-way sync into the mesh's memory.** A job reads this repository on a schedule and writes + the documents into the searchable store. It works with what exists today and needs nothing + built first. **Rejected**, because it creates a second copy of every document, and the failure + mode of a derived copy is the one this repository is least able to tolerate: *the copy that is + searched quietly stops matching the copy that is edited*, and the enforced one wins. A design + record that has silently diverged from the reasoning it claims to carry is worse than one that + cannot be found — the first misleads, the second merely fails. + +2. **Leave the signpost and close nothing.** Honest, free, and it keeps the gap visible. + **Rejected as an end state**, though it is what stands until the option below exists. It + answers only for a reader who already suspects these documents exist, which is precisely not + the reader the mesh's memory is designed for. + +3. **An agent reads this repository directly, and the search consults it.** Nothing is copied. + **Adopted.** + +## Decision + +**The design record is read where it is written.** Retrieval is an agent reading this repository, +not a copy living in a second store — and a search of the mesh's memory consults that agent, so +what it knows appears beside ordinary results rather than only when it is asked. + +Both halves are the decision. The first alone is merely a reader, and would leave this repository +reachable but not surfacing — the state measured above. **The second half is what discharges the +promise** that these documents are returned beside everything else. + +**There is no copy, and that is the point.** No sync, no schedule, no reconciliation, and nothing +that can drift, because there is only ever one of each document. It is also always current, +including for work that is not yet committed. + +**The direction of reading is one-way and stays that way.** The agent reads this repository and +answers from it. Nothing flows back: this repository is public, the mesh is not, and a return path +would be how installation-specific detail arrives into documents that must not carry it +([`README.md`](../README.md)). + +## Consequences + +**This repository stops being a fourth knowledge system, properly.** The original objection was +about adding a knowledge *system*. An agent with read access adds no store at all — which answers +the objection more completely than the indexing that was promised, rather than merely as well. + +**ADR 0019's promise is amended, not satisfied.** It said these documents would be *indexed*. They +will not be. They will be *read*, and the search will ask. The commitment that survives is the one +that mattered — that a searcher finds them without already suspecting they exist — and the +mechanism behind it is different from the one named. + +**It is gated on an agent that does not exist yet.** Until it does, the signpost is what stands, +and this repository is reachable rather than surfacing. That is a known and stated gap, not a +silent one — and the gap is now a build task with a decided shape rather than an open question. + +**The search must degrade honestly.** When the agent cannot be reached, a search has to say that +this material was not consulted. A result set that silently omits it looks identical to one where +nothing matched, and *silence and success must never look alike* +([ADR 0004](0004-a-node-and-how-it-joins.md)) — the rule this repository has now paid for twice. + +**A rule states how it is checked, and this one is checkable.** The check is the measurement that +produced this record: search the mesh's memory for a phrase that appears only in a design document +here, and require it back. That check fails today, deliberately, and passing it is what closes +`04-ISSUES/006`. + +## References + +- [`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md) — + the gap, the two measurements, and why closing it early was refused +- [ADR 0019](0019-how-this-repository-works.md) — the promise this amends +- [`README.md`](../README.md) — the objection, and the gap named where the claim used to sit diff --git a/02-DECISIONS/0026-every-decision-is-a-record.md b/02-DECISIONS/0026-every-decision-is-a-record.md deleted file mode 100644 index eb64b60..0000000 --- a/02-DECISIONS/0026-every-decision-is-a-record.md +++ /dev/null @@ -1,89 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 26. Every decision is a record; there is no ledger - -## Context - -HQ carried a decision ledger at its root: a chronological table of forty-one numbered -decisions, each with who decided and a pointer to where the reasoning lived. It was created -deliberately, to make decisions findable and to give a home to decisions too small to warrant -a document. - -By the time the decision records were back-filled -([ADR 0024](0024-the-numbering-is-the-flow.md)) the ledger had become three things at once, -and only one of them was still needed. - -Classified, its forty-one entries were: ten restating a record, eleven restating design -documents, fifteen describing how this repository works — with the reasoning in a README rather -than anywhere citable — three small rules with no home at all, and two superseded stubs. - -So the ledger was mostly a copy. Worse, it was a **hand-maintained index**, which -[ADR 0022](0022-status-lives-in-frontmatter.md) had just finished rejecting for the -decision-record index on the grounds that it had drifted after a single addition. Keeping one -copy of that pattern while removing another is not a position. - -It had also produced a naming collision that a directory listing makes plain: `DECISIONS.md` -beside `02-DECISIONS/`, holding different things. - -## Considered options - -1. **Keep the ledger.** Rejected. It duplicates the records, restates status, and is the exact - hand-maintained index this repository decided against elsewhere. -2. **Keep it, renamed, for small decisions only.** Rejected, and this is the option worth - arguing with — it is genuinely useful to record a decision without writing a document. But a - decision small enough to be one table row is almost always a **rule** rather than a - decision, and a rule belongs in [`how-we-build.md`](../00-META/how-we-build.md) where it is - enforced and where its reasoning is kept. That is where the three orphans went. -3. **Every decision is a record; nothing else.** Chosen. This is how the sibling HQ repository - for the PAPA platform works, and it has no ledger of any kind. - -## Decision - -**If a decision is worth recording, it is worth a record. If it is not worth a record, it is -not recorded.** - -`02-DECISIONS` holds every decision. There is no ledger, no index file, and no central status -of any kind. The chronological view — decisions in the order they were taken — is *generated* -from record frontmatter, which is what the ledger was actually for. - -Content that was only in the ledger was rehomed rather than dropped: - -| Was | Went to | -|---|---| -| Decisions about how this repository works | Records [0019](0019-hq-is-its-own-repository.md)–[0025](0025-hq-is-the-source-of-the-constitution.md) | -| Small rules with no record | [`how-we-build.md`](../00-META/how-we-build.md) — the package rule, and two already there | -| Lab decisions not stated in the design | [`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) | -| "Deliberately not decided" | The research effort and design document each question belongs to | -| Unowned observations | [`04-ISSUES`](../04-ISSUES/) ([ADR 0023](0023-issues-have-a-front-door.md)) | - -## Consequences - -- One place to look, and nothing to keep in sync. The collision between the ledger and the - record folder is gone. -- Structural parity with the sibling repository on decisions, which - [ADR 0024](0024-the-numbering-is-the-flow.md) deliberately broke on folder numbering. The - divergence is now exactly one thing, and it is the one thing that was argued for. -- **Writing a record is now the only way to record a decision, and a record is more work than - a table row.** The real risk is that a small decision goes unrecorded because nobody wanted - to write a document. The mitigation is that a small decision is usually a rule, and - `how-we-build.md` takes rules cheaply — but this is a cost, not a solved problem, and it is - the thing to watch. -- The chronological view now depends on the generator existing and being run. It did not - before. -- Two superseded ledger stubs had no record of their own. The position that documentation - lives inside the code repository is now recorded only as superseded context in - [ADR 0019](0019-hq-is-its-own-repository.md); the system-container position is explained in - [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md). Neither is lost. - -## References - -- The sibling PAPA HQ repository: root holds only agent instructions and a README; every - decision is a numbered record, and its graduation playbook has no path for an unrecorded - decision. -- [ADR 0022](0022-status-lives-in-frontmatter.md) — the hand-maintained-index argument this - applies consistently. diff --git a/02-DECISIONS/0026-the-mesh-has-a-session-of-its-own.md b/02-DECISIONS/0026-the-mesh-has-a-session-of-its-own.md new file mode 100644 index 0000000..7fbfb1d --- /dev/null +++ b/02-DECISIONS/0026-the-mesh-has-a-session-of-its-own.md @@ -0,0 +1,130 @@ +--- +topic: what runs on it +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0004-a-node-and-how-it-joins.md +--- + +# 26. The mesh has a session of its own, and it is the node session's mechanism + +## Context + +[ADR 0004](0004-a-node-and-how-it-joins.md) gives every node a session: one per node, permanent, +remembering across callers, its system prompt the node's engram, reachable over the broker like +everything else. **Any node can message any node**, and that is called the one part of the system +that is genuinely a mesh — symmetric, with no centre. + +**There is no way to address the mesh itself.** A question that spans machines — *what is running +across all of this*, *which nodes are behind*, *why is it built this way* — has to be put to some +node, which then asks the others. That works, and it makes a mesh-wide question **nobody's +question**: every node answers it as a foreigner, from a position where the whole is not in view. + +**Three things independently arrived at the same missing piece.** + +[ADR 0025](0025-the-design-record-is-read-not-copied.md), taken hours before this one, commits to +an agent that reads the design repository directly and answers into search. That agent has to +exist, run somewhere, and be askable — and nothing in the record says what it is or where it +lives. + +[`14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) records, as a gap deliberately +not half-built: *this worker uses that licence is a binding to an agent, not to a node* — and the +provisions model has no consumer identity other than a node. A session that must be assigned a +licence is exactly that consumer, and node sessions are already one. + +**And ADR 0004 never said how a session is set up.** It describes behaviour and stops: nothing +states how a session starts, where its context lives, how the engram reaches it, or how a message +off the broker becomes a prompt. There is no design document for it. That gap was invisible until +something had to be built *like* a node session, because describing a second instance of a +mechanism requires the mechanism to have been described once. + +## Considered Options + +1. **No mesh session; keep relaying through a node.** Costs nothing and works today. **Rejected.** + It leaves mesh-wide questions belonging to nobody, and it does not survive contact with + ADR 0025 — that agent still needs a home, so the thing gets built anyway, unnamed, as an + attachment to whichever node happened to host it. + +2. **A new kind of agent, built separately.** Purpose-built for the whole mesh. **Rejected.** It + would hold a session, a memory, a licence and broker plumbing — every one of which the node + session already has. Two implementations of one mechanism drift, and the vocabulary collision + that [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) exists to undo began exactly this + way: two things that were nearly the same, built twice, until neither word meant one thing. + +3. **The same mechanism, started in a different context.** **Adopted.** + +## Decision + +**The mesh has one session, addressed as the mesh, and it is a node session in every respect but +three.** + +| | | +|---|---| +| **the context it starts in** | the mesh's, not a machine's — this is the whole of what makes it different | +| **its engram** | the mesh's system prompt, as a node's engram is that node's | +| **its licence binding** | assigned in its own right, not inherited from the machine it runs on | + +Everything else is unchanged and deliberately so: it is permanent, it remembers, it is reachable +over the broker, it holds its own tools, and switched off it still answers *I am switched off* +rather than falling silent. + +**It runs on the node that holds the control plane** — not for convenience, but because that node +is already the one place excepted from *compromise of a node is compromise of that node* +(ADR 0004). An agent able to reach everything, placed anywhere else, creates a **second** such +place. Putting it where the authority already sits concentrates nothing new. + +**It is an addition to per-node messaging and never a replacement.** Every node remains directly +addressable. This is not a preference: ADR 0001 holds that losing the control plane costs *change, +not operation*, and a mesh whose only conversational surface lives on that node would lose the +ability to ask anything while every machine kept running perfectly. **The front door may not be +the single point.** + +**It is not an employee** ([ADR 0003](0003-agents-are-persistent-employees.md)). Nobody hires it, +it holds no task queue, it is never drained or reassigned. What it does with work that belongs +somewhere else is **dispatch it** — to node sessions, or to workers — which is what a node session +already does when asked something it does not have. + +**It is ADR 0025's reader.** The agent that reads the design repository and answers into search is +this session, not a second one. One agent, one memory, one place to reach; two would both need +that repository and would eventually disagree about what it says. + +**"One per node" is about address, not about process count.** ADR 0004's rule — *two and nothing +decides which replies* — forbids ambiguity in who answers when a **node** is addressed. The mesh +session answers when the **mesh** is addressed. The control-plane node therefore hosts two +sessions and no ambiguity, and stating this here is what stops it reading as a contradiction +later. + +## Consequences + +**The node session's setup must now be designed, and it never was.** This decision is expressed as +*the same as a node session, elsewhere*, which is only meaningful once that mechanism is written +down. The design document covering both is the immediate consequence of this record, not a +follow-up to it. + +**A consumer that is not a machine stops being deferrable.** The licence binding above is the gap +`14-model-access.md` names, and it now has two consumers rather than a hypothetical one. Until it +exists, a session's model access can only be expressed as *this module on this machine*, which +cannot say *this node's session uses the personal licence and the mesh's uses the company one* — +the thing the binding is for. + +**Symmetry is preserved, and it is worth being precise about why.** ADR 0004's claim is about what +a node can reach, and it is untouched: node-to-node messaging is unchanged, nothing is routed +through the mesh session, and it is a participant rather than a hop. What arrives is a +participant that happens to be the one a person usually addresses. + +**Availability degrades to inconvenience rather than to silence** — but only because of the +addition rule above. If that rule is ever relaxed, this consequence inverts, and it inverts +quietly: everything keeps working and nobody can ask about it. + +**The surface a person uses is not decided here.** That a board is a good place to talk to it is +likely and is not this record's business; the session is reachable over the broker like everything +else, and what puts a text box in front of it is a separate choice. + +## References + +- [ADR 0004](0004-a-node-and-how-it-joins.md) — the node session this extends +- [ADR 0025](0025-the-design-record-is-read-not-copied.md) — the reader this session is +- [ADR 0003](0003-agents-are-persistent-employees.md) — the vocabulary this is not +- [`03-DESIGN/01-to-be/14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) — *a + consumer that is not a machine*, the gap this makes concrete diff --git a/02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md b/02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md new file mode 100644 index 0000000..6e2bce4 --- /dev/null +++ b/02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md @@ -0,0 +1,109 @@ +--- +topic: what runs on it +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0009-modules-and-the-graph.md +--- + +# 27. A provision names what the consumer is coupled to, not the role it plays + +## Context + +Provisions are named after roles. The catalogue and every test fixture built so far say: + +``` +provides: database +requires: database +``` + +**Nothing distinguishes one engine from another.** A module requiring `database` is satisfied by +any module providing `database`, so a module written against PostgreSQL can be matched to a +provider of Microsoft SQL Server, resolve as satisfied, deploy, and fail on its first query. + +**The mesh runs several engines** — PostgreSQL, Microsoft SQL Server, MariaDB, and others behind +products that expose their own. This is not a hypothetical collision. + +**The failure is in the direction that hides.** Resolution *succeeds*. Nothing is refused, nothing +is logged, and the breakage surfaces later as an error inside an application, on a machine, with +nothing connecting it back to a match made elsewhere by something that thought it had done its +job. **A wrong answer delivered confidently costs more than a refusal**, and the whole point of +refusing on ambiguity ([ADR 0009](0009-modules-and-the-graph.md)) was to not do this. + +**How it got in:** every test written for the resolver had exactly one provider of each name, so no +mismatch was expressible and none was caught. The fixtures agreed with the design. That is the same +fault as [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)'s +imagined output and [`019`](../04-ISSUES/019-a-comment-asserting-a-fact-about-a-machine/00-report.md)'s +unchecked comment, at the level of a name rather than a line. + +## Considered Options + +1. **Keep role names; let the operator assign correctly.** The mesh would refuse ambiguity when two + providers exist, so a person picks. **Rejected.** It makes correctness depend on somebody + knowing that the module they are assigning speaks a particular dialect — which is exactly the + knowledge the provisioning model exists to remove. And with one provider of each name, nothing + is ambiguous and nothing is asked. + +2. **A role name plus a `flavour:` or `engine:` qualifier**, matched as a second field. + **Rejected.** Two fields that must agree is a constraint the resolver has to enforce and a + manifest author has to remember, to express something one field already can. The name is the + contract; splitting it invites a requirement that names a role and forgets the qualifier, which + then matches everything again. + +3. **The name says what the consumer is coupled to.** **Adopted.** + +## Decision + +**A provision is named for the thing a consumer's code is written against.** + +``` +provides: postgres-database +requires: postgres-database +``` + +**The test is whether the consumer can tell the difference.** If swapping the provider would break +the consumer, the name must say which provider — because a match that breaks the consumer is not a +match. If the consumer genuinely cannot tell, a role name is correct and better. + +| provision | | why | +|---|---|---| +| `postgres-database`, `mssql-database` | **specific** | applications are written against a dialect; a swap breaks them | +| `route` | **role** | the consumer wants its name reachable and does not care what proxies it | +| `resolver` | **role** | the consumer wants names to resolve | +| `artifact-store` | **role** | the consumer fetches by digest over a protocol, and nothing else | + +**`database` is not a provision and may not be provided.** There is no context in which an +application talks to a generic database: it talks to PostgreSQL or it talks to SQL Server. A name +that cannot be true of any real consumer should not be expressible. + +**This is about coupling, not about products.** Two providers of `postgres-database` — a container +on this node and a managed instance elsewhere — are interchangeable and *should* both match. What +may not be interchangeable is what the consumer's queries are written in. + +## Consequences + +**Every manifest that names a database changes.** Doing this now costs a rename across a handful of +examples. Doing it after modules are migrated costs it across all of them, plus every deployment +that resolved against the old name. + +**Wrong requirements now fail loudly, and at the right moment.** A module requiring +`postgres-database` where only `mssql-database` is provided is unsatisfiable, so it is **refused at +resolution** with both names visible — rather than deployed and broken later. This is the property +that was lost, restored. + +**Generic role names are still right, and the rule says when.** This does not push specificity +everywhere; it puts it exactly where a consumer is coupled. Naming `route` after a particular proxy +would be the same error in the other direction, and would prevent a swap that genuinely changes +nothing. + +**It is checked, not merely stated** ([`00-META/how-we-build.md`](../00-META/how-we-build.md) §5). +A manifest providing a name known to be engine-generic is refused, naming what to say instead. +Without that, this record is a convention, and a convention is what the previous naming was. + +## References + +- [ADR 0009](0009-modules-and-the-graph.md) — provisions, and refusing on ambiguity +- [`03-DESIGN/01-to-be/07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md) — *the + provisioning model uses databases, roles and schemas as PostgreSQL means them*, which is this + record's point made about the substrate before it was made about modules diff --git a/02-DECISIONS/0027-the-product-is-novox-mesh.md b/02-DECISIONS/0027-the-product-is-novox-mesh.md deleted file mode 100644 index 192b7a3..0000000 --- a/02-DECISIONS/0027-the-product-is-novox-mesh.md +++ /dev/null @@ -1,112 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 27. The product is Novox Mesh; Nox is an identity, not a second system - -## Context - -The name `HAL` was never chosen. This began as a dotfiles repository, the first commits in -February 2026 adopt dotfiles and per-node overrides, and the name arrived with the code — as -recorded in -[`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md), -most of the current shape is inherited from that origin rather than designed for a mesh. The -name is part of the inheritance. - -Three things make it worth changing rather than living with. - -**It is borrowed, and borrowed badly.** HAL is the canonical *untrustworthy* machine -intelligence. For infrastructure whose entire proposition is that it manages your machines, -heals itself, and is trusted with credentials, that is an unhelpful flag to fly, and it is not -a name anyone owns. - -**There is a name available that is owned.** The company is Novox. A product of Novox should -carry that lineage rather than a film reference. - -**A platform and a persona are different things, and one name was doing both.** `HAL` named the -mesh *and*, implicitly, the thing an operator talks to. Those are separate concerns — the -platform is what runs; the persona is who answers. - -## Considered options - -1. **Keep `HAL`.** Rejected. Every reason to keep it is sunk cost, and the sunk cost is at its - smallest today. -2. **Rename everything to a single new name covering platform and persona.** Rejected: it - repeats the conflation that made `HAL` ambiguous. -3. **Separate the two: a product name and an identity.** Chosen. - -## Decision - -**The product is `Novox Mesh`**, shortened to `mesh` in internal use — repository names, the -module namespace, environment variables, paths. - -**`Nox` is an identity of Novox**, and specifically an **agent identity within the mesh's own -model** — a named participant, exactly as -[ADR 0012](0012-agents-are-persistent-employees.md) defines one. Not a separate product, not a -separate runtime, not a privileged path. - -**Nox is the agent of the mesh, not of a node.** This is the part that carries weight: - -- **Every node keeps its own identity.** That already exists and stays — a node is a named - participant with its own character, and addressing one directly remains possible and normal. -- **Nox is scoped to the whole mesh.** It is what the mesh is called when the mesh itself - speaks, rather than one machine within it. -- **Nox addresses node identities.** Asking Nox for something that lives on one node is Nox - talking to that node, not a human choosing a machine. -- **A human mostly talks to Nox.** It is the front door. - -That last point makes Nox the concrete form of the vision in -[`00-META/mission.md`](../00-META/mission.md): *an agent states an intent and the mesh carries -it out — no console to open, no runbook to follow, no remembering which node holds which -thing.* Nox is who that intent is stated to. The mission described the behaviour; this names -the thing that has it. - -Nox holds no private channel. Whatever it can do, it does through the same surfaces every other -agent uses — which is not a naming detail: a persona with its own path would be the one part of -the mesh with no human checkpoint, and the skeleton already rules that out. - -`HAL` is retired. - -**Timing is the substance of this decision, not an aside.** The skeleton in -[research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) is not built. Renaming -before it exists costs a search and replace across research documents. Renaming after costs the -same class of migration as everything else this repository is trying to avoid, and would -therefore not happen. - -## Consequences - -- **The as-is layer keeps `HAL`.** It describes what runs, and what runs is called HAL. The - to-be layer uses `mesh`. The rename is part of the migration, and the two-layer split - ([ADR 0020](0020-design-is-written-in-two-layers.md)) is what makes holding both names - coherent rather than confusing. -- **Records 0001–0026 keep `HAL`.** They are immutable and they say what was decided when it - was decided. No record is edited for a name. -- Tier 2 cannot be `mesh-mesh`. The control plane is **`mesh-control`**; `mesh-broker` was - rejected because the substrate already contains a message broker. -- The namespace, environment variable prefix, service names and on-disk paths all change. In - the existing system that is a migration and is not attempted here. -- **`mesh` is a generic word**, and it already means something specific in infrastructure — a - service mesh is a different thing. Recorded as a known trade rather than an oversight: the - full name `Novox Mesh` is distinctive, and the short form is internal. -- The persona has a name before it has behaviour. That is the right order — it is an identity in - a system that already has a model of identities, so it needs no new machinery to exist. -- **Except in one respect, and it is a real gap.** - [ADR 0012](0012-agents-are-persistent-employees.md) binds every agent to a home node, one to - one, with a workspace on that machine. A mesh-scoped agent has no home node by definition, so - the model does not currently have a shape for Nox. Extending it — an agent whose scope is the - mesh rather than a machine — is a decision of its own and is not taken here. -- Two levels of identity now exist where there was one: the node, and the mesh. The distinction - has to stay visible in every surface, or "ask Nox" and "ask a node" collapse into each other - and it stops being clear who is answering. - -## References - -- [ADR 0012](0012-agents-are-persistent-employees.md) — what an identity is in this system, and - why `Nox` needs no separate mechanism. -- [ADR 0020](0020-design-is-written-in-two-layers.md) — why the as-is and to-be layers can - legitimately use different names for the same system. -- The dotfiles origin, and the naming inheritance it explains: - [`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md). diff --git a/02-DECISIONS/0028-hq-is-company-scoped.md b/02-DECISIONS/0028-hq-is-company-scoped.md deleted file mode 100644 index 24d6b06..0000000 --- a/02-DECISIONS/0028-hq-is-company-scoped.md +++ /dev/null @@ -1,89 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 28. HQ is company-scoped; the mesh is its first product - -## Context - -This repository was `hal-hq` — one product's headquarters, named for the product. Then the -product was renamed ([ADR 0027](0027-the-product-is-novox-mesh.md)), which forced the question -of what the repository is actually the headquarters *of*. - -Two facts settled it, and both were checked rather than assumed. - -**Novox already delivers other things.** The company's forge organisation holds live projects -beside the mesh, and they are registered as build sources — meaning the mesh already builds and -deploys them. They are not hypothetical future products; they exist and ship today. - -**They are tenants, not peers.** They run *on* the mesh. Every one of them is developed, -delivered and hosted by it. So the mesh is not one product among several — it is the ground the -others stand on. - -That distinction decides the scope. If the mesh were a product beside others, a per-product HQ -would be right. Because it is the substrate the company operates on, a decision about the mesh -is a decision about how the company works. - -## Considered options - -1. **`mesh-hq` — one HQ per product.** The safe choice, and the reversible one: a second - product creates its own HQ and shared practice graduates upward later. Rejected, knowingly, - because it models the mesh as a peer of things that are actually its tenants. -2. **A company HQ *and* a product HQ, from the start.** Rejected as ceremony — two repositories - for one operator, and the constitution's own YAGNI rule says not to. -3. **One company-scoped HQ, `novox/hq`, with the mesh as its first product.** Chosen. - -## Decision - -The repository is **`novox/hq`** — Novox's headquarters, not the mesh's. - -It holds the reasoning behind what Novox builds. Today almost all of that is the mesh, because -the mesh is what Novox is building. That is a fact about the present, not a definition of the -repository. - -**The scope of each document is fixed now, so the eventual split is mechanical rather than -archaeological:** - -| Scope | Documents | Moves if products separate? | -|---|---|---| -| **Company** | [`how-we-build.md`](../00-META/how-we-build.md), [`process/`](../00-META/process/), [`repos.md`](../00-META/repos.md), this record and [0019](0019-hq-is-its-own-repository.md)–[0027](0027-the-product-is-novox-mesh.md) | No — they stay at the top | -| **Product (mesh)** | [`mission.md`](../00-META/mission.md), [`context.md`](../00-META/context.md), [`effect.md`](../00-META/effect.md), `01-RESEARCH`, `03-DESIGN`, `04-ISSUES`, records 0001–0018 | Yes — into a product section | - -The folders are **not** restructured now. One product's content under a company name is -correct while there is one product's worth of it, and nesting before there is anything to nest -is the ceremony option 2 was rejected for. - -## Consequences - -- Engineering practice has a home that does not belong to the mesh. `how-we-build.md` — never - write to production directly, migrations for schema changes, runtime evidence for behavioural - criteria — is true of any Novox project, and its being in a mesh repository was always a - slight mislabelling. -- The constitution derived from it ([ADR 0025](0025-hq-is-the-source-of-the-constitution.md)) - can legitimately govern work outside the mesh. Under a product HQ it could not have, without - either duplicating or reaching across repositories. -- **The bet is not entirely forward-looking, and that is worth being honest about.** Novox - already has work that is *not* a mesh tenant — client engagements and at least one product - that is developed outside it. So the company genuinely has a scope wider than the mesh - **today**, which strengthens the case for a company HQ and simultaneously means the split in - the table above is closer than "some day". The table is not a precaution; it is a plan whose - trigger already half-exists. -- What has *not* happened yet is any of that work needing the constitution. That is the actual - trigger ([ADR 0025](0025-hq-is-the-source-of-the-constitution.md)): the moment something - outside the mesh must be governed by the same rules, product-level content moves down a level - and this repository becomes what its name already claims. -- A new repository was created rather than the old one transferred, because the forge predates - the transfer API. The original was verified to contain nothing the new one lacks — every ref - an ancestor, no tags, issues, pull requests, releases or wiki content — and then removed. -- The mesh's own documents now live one conceptual level below the repository they are in. A - reader arriving at `01-RESEARCH` should understand it as the mesh's research, not Novox's. - Nothing in the folder names says so, and that is the cost of not restructuring. - -## References - -- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name that forced the question. -- [ADR 0019](0019-hq-is-its-own-repository.md) — why HQ is a repository at all. Unchanged; only - its scope moves. diff --git a/02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md b/02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md new file mode 100644 index 0000000..0b6bc06 --- /dev/null +++ b/02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md @@ -0,0 +1,118 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md +--- + +# 28. The substrate supplies the control plane and nothing else + +*Corrects one row of [ADR 0006](0006-the-substrate-and-the-control-plane.md) and makes explicit +something it left unsaid. The rest of that record stands.* + +## Context + +ADR 0006 defines the substrate by a circularity: **what the control plane needs in order to run, +and cannot ask itself for, because it is not running yet.** Two questions, and both must be +answered *yes* for something to be substrate. + +Its membership table admits the object store on this line: + +| role | product | | +|---|---|---| +| object store | **MinIO** | it cannot grant itself a bucket | + +**That answers the second question and assumes the first.** It is true that a control plane cannot +grant itself a bucket. Nothing establishes that it needs one. + +**It does not.** Verified 2026-08-31 against `mesh-control`: no S3 client, no bucket, no object +storage of any kind outside comments. Artifacts reach nodes as content-addressed blobs in the OCI +registry, and the code records the decision and its reasoning: + +> One store, and it is the registry the bootstrap already pulls from. An OCI registry is a +> content-addressed blob store that happens to also understand images… The alternative considered +> was a second store beside it — S3-shaped, buckets, signed URLs. It is the right answer for +> objects that are *mutable*, or need per-reader access, or are not build output. None of that +> describes a digest-pinned archive, and standing up a second service to hold one kind of +> immutable blob means two things to run, two things to back up and two ways for an artifact to be +> missing. + +**The row is inherited from the system being replaced**, where an object store distributed module +tarballs. Here nothing does, and the row was never re-tested against the definition it sits under. + +**A second thing ADR 0006 never says:** whether a substrate service and a module of the same +product are the same instance. It says the substrate is *not the control plane* and *not a place +for logic*, and stops. The question is not idle — an application wanting a database, on a mesh +whose substrate is already running PostgreSQL, has an obvious wrong answer available. + +## Considered Options + +1. **Leave the object store as substrate, unused.** Harmless-looking. **Rejected.** A membership + list that includes something nothing needs is a list that has stopped being derived from its + test, and the next member is admitted by precedent instead of argument. It also mandates that + every mesh run a service no mesh uses. + +2. **Applications share the substrate's instances.** One PostgreSQL, one of everything. + **Rejected**, below. + +3. **The substrate is exactly what the control plane consumes; everything else is a module.** + **Adopted.** + +## Decision + +**The object store is not substrate.** It fails the first half of the test: the control plane does +not need one. An object store is an ordinary module, required through the module graph like +anything else, and a module wanting one depends on a module providing one. + +**The substrate has four members, not five**: a relational store, a message bus, an image registry, +and conditionally an identity provider. The registry stays — the control plane genuinely cannot +deliver an artifact without somewhere to put it. + +**A substrate service and a module of the same product are different instances, and are not +shared.** The mesh's own PostgreSQL and a PostgreSQL a workload was given are two servers, two +containers, two lifecycles. + +Three reasons, and the first is the one that matters: + +**The substrate is not in the module graph.** It is raised from the pinned bundle the host carries, +before any mesh exists to declare it. A workload depending on it would depend on something the +graph cannot see, cannot rotate a credential for, and cannot move — which is every property the +provisioning model exists to provide. + +**It would put workload data in the control plane's own store.** The mesh's contexts own their +stores exclusively ([ADR 0008](0008-a-context-owns-its-store.md)). An application sharing that +server can exhaust it, lock it, or fill its disk, and the failure is the control plane going down +— which is the one failure that makes every other one harder to fix. + +**They are bounded differently.** The substrate is sized, backed up and upgraded as part of +bootstrapping a mesh. A workload's database follows the workload — moved with it, destroyed with +it, restored with it. + +## Consequences + +**Migrating an object store is ordinary module work**, not substrate work. It was previously going +to be done as part of completing the substrate, which would have been the wrong shape and would +have coupled every mesh to a service the mesh does not use. + +**A mesh with no workload needing one runs no object store at all.** That is the correct outcome +and was not previously available. + +**Two PostgreSQL containers on a node that hosts both is expected**, not duplication to be +optimised away. Anyone tidying them together should find this record first. + +**"Substrate by role and ordinary by delivery" loses one of its two members.** ADR 0006 uses that +phrase of the object store and the registry — things that are substrate but provisioned once a +control plane exists. It now describes the registry alone. + +**The definition is applied, not just stated.** Both halves of the circularity test are asked of +each member, and *cannot grant itself one* is not sufficient on its own — it is true of almost any +service, which is what made it possible to admit a member on that half alone. + +## References + +- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the definition, and the table this + corrects one row of +- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store exclusively +- `mesh-control internal/builder/registry.go` — where artifacts go, and why not S3 diff --git a/02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md b/02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md new file mode 100644 index 0000000..94acc9b --- /dev/null +++ b/02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md @@ -0,0 +1,112 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0005-the-node-host.md +--- + +# 29. A network is a shape, because an action cannot be undone + +## Context + +**A module of several containers has no way to let them reach each other by name.** A container +declaration carries a `network` field, and it only ever *joins* one that already exists — it was +added so the control plane could reach the store and the broker on the machine it was raised on. +Nothing in the vocabulary **creates** a network. + +Without one, containers on a machine share the runtime's default bridge, which gives addresses and +no name resolution between them. So a module that is several containers can only be written by +publishing ports onto the machine and pointing its own parts at the host — which puts a module's +private wiring on the machine's own address space, where anything else on the machine can reach it +and any other module can collide with it. + +**This is the gap a mail system meets and nothing else so far does** +([`00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) 3.3). It is being taken now +rather than then, because 3.3 is the task most likely to send work back into the declaration +language and the least useful place to discover it. + +**Adding a shape is not a small change, and the host says so** — the vocabulary is asserted +against a stated number, with the reason written into the failure: *every addition widens what a +compromised control plane can express, so a change here is a decision.* The host applies what it +is told; the only bound on a hostile control plane is what the language can say +([ADR 0004](0004-a-node-and-how-it-joins.md)). + +## Considered Options + +1. **An `action` that creates the network.** The vocabulary already has one, the bundle already + uses seven of them, and `docker network create` with a `verify` is exactly the shape an action + takes. Nothing would need adding. **Rejected**, on removal: + + > An action has no footprint the host can undo — it ran, and whatever it did belongs to + > whatever it acted on. + + A network made this way **leaks when the module is unassigned**, and the mesh cannot tell: the + record says an action ran, and there is nothing to reverse. Unassigning a module would leave a + network behind on every machine it was ever on, and the only way to find them would be to go + and look. *A resource the mesh can create and never clean up is one it should not create.* + + There is a second reason, and it is the one that generalises: an action is opaque. **The mesh + cannot tell what an action did**, so a network created by one is not a thing the mesh knows + about — it cannot be reported, counted, or reasoned about, and a module could not require one. + +2. **Publish ports on the machine instead.** No new shape, and it works today. **Rejected.** It + makes a module's internal wiring part of the machine's address space: two modules that each + want a database on a fixed port collide, and anything else on the machine can reach what was + meant to be private. It also makes the module's manifest depend on what else is installed, + which is the thing provisioning exists to remove. + +3. **`network` as a ninth shape.** **Adopted.** + +## Decision + +**`network` joins the vocabulary, and the vocabulary is nine shapes.** + +``` +{"id": "internal", "type": "network", "name": "mail"} +``` + +**A name and nothing else.** Not a driver, a subnet, an address range or a gateway: every one of +those is a thing a module would have to know about the machine it lands on, and a module that +names a subnet is a module that collides with whatever else chose the same one. The runtime picks; +the mesh names. + +**It is created if absent and removed when no longer declared** — an ordinary shape, with the same +lifecycle as a directory. That is the whole reason it is a shape. + +**Declared before the containers that join it.** Resources are applied in the order the module +wrote them, and orphans are removed in **reverse** — so a network written first is created first +and removed last, after the containers attached to it are gone. This is not a new rule; it is the +existing one, and it happens to be exactly right here. A network written *after* its containers +would fail to remove while they still hold it, and that failure is reported rather than silent. + +**What it does not do:** it does not reach across machines. A network is one machine's, like +everything else the host applies. Modules on different machines reach each other over the private +network the mesh already provides ([ADR 0007](0007-connectivity.md)), and a shape that tried to +span machines would be a second overlay with a worse contract. + +## Consequences + +**The vocabulary is nine, and the count moves with a record.** The test that asserts it names this +one, so the next person to change it finds the argument rather than a number to edit. + +**A compromised control plane can now create and destroy networks on a machine.** Stated plainly +because that is the cost, and the bound is the point: it can create a named network and remove +one, and it can do neither to anything it did not declare. It cannot inspect, attach to, or +reroute what is already there — those would be different shapes, and are not being added. + +**A multi-container module becomes expressible**, which unblocks 3.3 and, less obviously, makes +several smaller modules simpler: anything that is a service plus a sidecar currently has to +publish a port to talk to itself. + +**Nothing is required to use it.** A module of one container declares no network and joins none, +exactly as now. The substrate keeps using `host`, which is a runtime-provided network and not one +the mesh creates. + +## References + +- [ADR 0005](0005-the-node-host.md) — the host's vocabulary, and why each shape is a decision +- [ADR 0004](0004-a-node-and-how-it-joins.md) — what may be pushed is bounded by form, not by trust +- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — 1.3, + and the mail system at 3.3 that this is for diff --git a/02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md b/02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md deleted file mode 100644 index 0ca6c40..0000000 --- a/02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md +++ /dev/null @@ -1,101 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false -extends: 0016-a-lab-node-is-a-virtual-machine.md ---- - -# 29. The lab's first scenario has no pipeline, and the lab comes first - -## Context - -[The lab design](../03-DESIGN/01-to-be/01-end-to-end-testing.md) opens with *"what is under -test is a module; the mesh is the harness"*, and everything follows from that: a scenario has -its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict -*is* a pipeline result. - -That is the right design for testing a module against the mesh that exists. It is unusable for -the thing now being built. - -**The new mesh has no coordinator.** Tier 0 is a host binary and tier 1 is a pinned bundle -([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)). A scenario that -requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because -all four are tier 2 and do not exist yet. - -**And the sequencing was written backwards.** [Research 009](../01-RESEARCH/009-migration/00-overview.md) -placed the lab at phase B, as verification of tiers already built. But tier 0 is the component -that takes over a machine's packages, services and network — it cannot be developed against a -machine anyone needs. It needs somewhere disposable to exist **before** it is written, not -after. - -## Considered options - -1. **Develop tiers 0 and 1 against a real machine; add the lab afterwards.** Rejected twice - over. Developing something that reformats a machine, against a machine that is in use, is - how a machine is lost. And it would leave the bootstrap path exercised only when performed - for real — which is precisely the property that makes the current first-node script the - least-tested code in the system. -2. **Build the full lab first.** Impossible, not merely unwise: the full scenario needs a - coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the - tiers it is meant to test. -3. **Two scenario classes, the smaller one first, the larger a superset.** Chosen. - -## Decision - -The lab has **two scenario classes**, and the first has no pipeline in it at all. - -| | **Bootstrap scenario** | **Full scenario** | -|---|---|---| -| Contains | one or more virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules | -| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify | -| Exercises | tiers 0 and 1 | tiers 2 and above, and modules | -| Exists to | develop the mesh | test what runs on it | - -The bootstrap scenario is a **strict subset** of the full one — the same virtualisation, the -same networking, the same scenario lifecycle, simply stopping before a control plane exists. -Nothing forks, which is the same rule the existing design already holds itself to. - -**The lab is built first**, ahead of tier 0, and [research 009](../01-RESEARCH/009-migration/00-overview.md) -is resequenced accordingly. It is the environment everything else is developed inside. - -Of the runner's two candidate jobs, this settles their order: **scenario lifecycle is needed -immediately** — something must materialise, snapshot and destroy a mesh before anything else -can be written. **Assertion execution comes later**, with the full scenario, because a -bootstrap scenario's assertions are about the state a single host reconciled and are small -enough to state directly. - -## Consequences - -- **The hardest path to test becomes the one exercised most.** Raising a node from nothing is - currently a script that runs when a node is created and is otherwise never touched. Under - this decision it is the inner development loop for every change to tiers 0 and 1. -- The first thing built is small: virtualisation, a network, a way to place a binary, and a way - to snapshot and reset. No forge, no coordinator, no pipeline, no modules. -- The full scenario becomes reachable by *addition* rather than by rework, because it differs - only in what is placed inside the machines. -- The lab acquires a second audience. It was designed for a module author and now also serves - whoever is building the mesh itself — which is the same "one runner, two callers" argument - the design already makes, extended one step. -- **A stale claim in the design is corrected.** It argues that scenarios are *"affordable with - system containers and would not be with virtual machines — the unit choice is what makes the - gate possible at all."* [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) superseded that: - a lab node is a virtual machine, and the scale argument for system containers was found to - have been invented rather than required. The design text did not follow the decision. It does - now. -- The lab's home is `novox/mesh-lab`, recorded in - [ADR 0030](0030-the-repository-structure.md) — written after this record, because this one - needed a repository that no decision had yet named. -- The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from - how a real node is raised, it certifies something that does not happen. That is the same - hazard the existing design names for the full scenario, and the same answer applies — - nothing new drives it, and what runs is the real thing. - -## References - -- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — a lab node is a virtual machine. -- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the - observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned - bundle, which is also exactly the bootstrap path. -- [Research 009](../01-RESEARCH/009-migration/00-overview.md) — the migration sequence this - reorders. diff --git a/02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md b/02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md new file mode 100644 index 0000000..a792ee8 --- /dev/null +++ b/02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md @@ -0,0 +1,96 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0005-the-node-host.md +--- + +# 30. Data outlives the mesh that declared it + +## Context + +**The conversion runs on live services holding real data**, and starts on the node that holds all +of it. Identity, mail, everything. The requirement stated plainly: a data directory may be +*moved*, and may never be *lost*. + +**The host deleted them.** A directory that stopped being declared was an orphan, and an orphan +directory was removed with `os.RemoveAll` — everything under it — while the report said +`removed`. A module unassigned took its database's files with it, and nothing anywhere said what +had been in there. + +Reproduced before it was fixed: assign a module, let a service write into its directory, unassign +the module, and the file is gone. + +**A directory stops being declared for ordinary reasons**, which is what makes this sharp rather +than theoretical. A module unassigned from a node. A manifest edited to move a data folder — the +exact operation the conversion needs. A resource renamed. A typo. **Every one of those is a normal +day's work, and every one of them was destructive.** + +**The removal order was already right, and that is what makes a fix possible.** Everything the +mesh puts inside a directory is itself a declared resource, and orphans are removed in reverse +declaration order — so by the time a directory is reached, what the mesh wrote there is already +gone. Anything still present was put there by something else. + +## Considered Options + +1. **A `keep` flag on the directory.** A module declares which of its directories hold data, and + the host leaves those. **Rejected.** It is safe only when somebody remembered, and the failure + of forgetting is total and silent. A rule that protects data only when it was asked to is not + a rule about data, it is a rule about attentiveness — and this is the one place in the system + where being wrong does not recover. + +2. **Never remove a directory.** Simple and unarguably safe. **Rejected**, narrowly: every module + ever assigned would leave its directories behind for ever, and a machine that accumulates + things nobody can account for is one where nobody can tell what is still in use. The clean-up + that is genuinely the mesh's is worth keeping. + +3. **Remove a directory only when it is empty.** **Adopted.** + +## Decision + +**A directory that still holds anything is kept, and the mesh says so.** An empty one is removed. + +**This is the host's existing line applied to the one shape where getting it wrong is +unrecoverable** — *it removes what it made and leaves what it merely configured* +([ADR 0005](0005-the-node-host.md)). An empty directory is what the host made. A full one is not. + +**No flag, no declaration, nothing to remember.** Emptiness is the test, and it is derived from +the removal order rather than asserted: the mesh's own contents are gone by then, so what remains +is by definition something nobody declared. + +**It is reported, not silent.** The outcome is `kept`, naming how many items are inside and saying +they are for a person to deal with. A directory quietly left behind is how a machine accumulates +things nobody can account for — which is the objection to option 2, and it is answered by saying +so rather than by deleting. + +**Files are unchanged.** A declared file is the mesh's own — it wrote it, it owns it, and losing a +configuration file is not the failure this is about. The distinction is deliberate: **directories +hold what other things produced; files are what the mesh itself put there.** + +## Consequences + +**Moving a data directory is now safe by default.** The manifest changes, the old path stops being +declared, and the data stays where it is until somebody has looked at it. That was the operation +most likely to destroy something during the conversion, and it is now the operation that does the +least. + +**Unassigning a module leaves its data.** Correct, and it means unassignment is no longer a way to +clean up — removing data is a person's act, done knowingly. Given what unassignment did before, +that is the trade being made and it is the right way round. + +**A machine can accumulate directories nobody removed.** Accepted, and mitigated by saying so +every time rather than by a periodic sweep. A sweep would be the deletion this record exists to +prevent, on a timer, with nobody watching. + +**It is not a backup, and must not be mistaken for one.** This stops the mesh destroying data. It +does nothing about a disk, a mistaken `rm`, or a service corrupting its own store. The conversion +still needs backups taken and **restored** before anything is moved — a backup nobody has restored +is a belief, not a copy. + +## References + +- [ADR 0005](0005-the-node-host.md) — the host removes what it made +- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — the + conversion this was found by planning diff --git a/02-DECISIONS/0030-the-repository-structure.md b/02-DECISIONS/0030-the-repository-structure.md deleted file mode 100644 index 4dca3ae..0000000 --- a/02-DECISIONS/0030-the-repository-structure.md +++ /dev/null @@ -1,99 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 30. The repository structure, and the rule that names them - -## Context - -The tiers are settled ([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)) -and the product is named ([ADR 0027](0027-the-product-is-novox-mesh.md)), but the repositories -themselves were only ever sketched in research. Two consequences had already appeared. - -[ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) makes the lab phase 0 of the entire -migration and **could not say where it lives**, because no record named a repository. - -And the research contradicted an accepted record: it listed `mesh-hq` for this repository, while -[ADR 0028](0028-hq-is-company-scoped.md) had decided `novox/hq` and explicitly rejected that -name. A design resting on research is a design resting on something that can change without a -decision. - -There is also an implied naming rule that has never been written down. ADR 0027 says -repository names take `mesh`; ADR 0028 gives this repository no prefix at all. Both are right, -for a reason neither states. - -## Considered options - -Only the naming rule had genuine alternatives; the tier repositories follow from the tiers. - -1. **No prefix — `novox/host`, `novox/control`.** The organisation already says Novox, so the - prefix reads as stutter. Rejected once it was established that Novox delivers more than the - mesh: with several products the prefix is not stutter, it is the product namespace doing - real work, and the forge has no nested groups to do it instead. -2. **An organisation per product — `novox-mesh/host`.** Puts the product boundary where the - forge's only real grouping primitive lives, so permissions and teams attach to it. Rejected - for now as premature: no per-product access boundary exists yet, and it costs `novox-` - repeated across every organisation. -3. **Product-prefixed repositories in the company organisation.** Chosen. - -## Decision - -**The naming rule:** a repository that belongs to a product carries that product's prefix. A -repository that is company-scoped does not. - -That is why this one is `hq` and the mesh's are `mesh-*`. Both records were already correct; -the rule connecting them is stated here. - -**The repositories:** - -| Repository | Tier | Holds | -|---|---|---| -| `novox/mesh-host` | 0 | the node host — the one binary installed by hand | -| `novox/mesh-substrate` | 1 | the four pinned services, as declarations | -| `novox/mesh-control` | 2 | the control plane and its contexts | -| `novox/mesh-surfaces` | 3 | tools, web, cli — thin, no logic | -| `novox/mesh-sdk` | — | contracts shared across tiers: types, not behaviour | -| `novox/mesh-lab` | — | the lab: scenario lifecycle, networking, placement | -| `novox/hq` | — | this repository. Company-scoped ([ADR 0028](0028-hq-is-company-scoped.md)) | - -**The lab is its own repository.** Its lifecycle differs from everything else in the list: it -is never shipped to a node, it outlives any single tier, and it drives virtualisation on a -workstation — which nothing else in the mesh does. Putting it inside the host would couple -development tooling to a shipped component; putting it inside the control plane would make the -bootstrap scenario depend on a tier that does not exist when it is needed. - -**Tier 4 is deliberately not decided here.** Whether the catalogue is one repository, one per -domain, or one per application remains open from -[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) and is blocked on -[research 005](../01-RESEARCH/005-domain-grouping/00-overview.md): how many repositories hold -domains cannot be answered before what the domains are. Recording the gap is the point — -`mesh-catalog` appears in the research sketch and is **not** decided by this record. - -## Consequences - -- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) can name its target. Phase 0 has - a home, which was the immediate blocker. -- The research sketch stops being load-bearing. It remains what it is — a sketch — and the - design layer can now cite a record instead. -- **Seven repositories where there is currently one**, for a mesh that today lives in a single - monorepo. That is the cost, and it is not small: seven release cadences, seven sets of - dependencies, and cross-repository changes that were previously one commit. The offsetting - argument is the tier rule — a boundary that only points downward is enforceable across - repositories and merely conventional inside one. -- The prefix will read as redundant for as long as the mesh is the only product with - repositories. That is accepted deliberately: the alternative is renaming everything at the - moment a second product appears, which is the class of migration this project is trying to - stop performing. -- Nothing is created yet. This records what the repositories *are*; creating them is part of - phase 0 and after. - -## References - -- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name the prefix comes from. -- [ADR 0028](0028-hq-is-company-scoped.md) — why this repository has no prefix. -- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lab, and why it is first. -- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the - sketch this supersedes as a source. diff --git a/02-DECISIONS/0031-the-control-plane-authenticates-nobody.md b/02-DECISIONS/0031-the-control-plane-authenticates-nobody.md new file mode 100644 index 0000000..850bf76 --- /dev/null +++ b/02-DECISIONS/0031-the-control-plane-authenticates-nobody.md @@ -0,0 +1,72 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md +--- + +# 31. The control plane authenticates nobody, so identity is a module + +## Context + +[ADR 0006](0006-the-substrate-and-the-control-plane.md) left one member of the substrate +conditional, and said exactly why: + +| role | product | | +|---|---|---| +| identity provider | — | **conditional**: substrate only if the control plane delegates authentication, which is undecided | + +[`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md) carried it as an open question — +*whether identity is the fifth* — noting it followed from a decision nobody had taken. + +**The decision is taken: the control plane does not delegate authentication.** There is no mesh +identity provider. + +**Nothing in the mesh's own machinery ever needed one.** A node proves itself with a keypair it +generated, over a broker account issued at enrolment +([ADR 0004](0004-a-node-and-how-it-joins.md)). Declarations are verified by signature. None of +that touches an identity provider, and the conditional was never about machines — it was only ever +about whether a *person* signing in to a mesh surface would be authenticated by something else. + +## Decision + +**Identity is a module**, like the mail system and the forge. It runs *on* the mesh, not *of* it +([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)) — a provider other modules require, +which is the ordinary shape and needs nothing new to express. + +**So the substrate is three, and no longer conditional**: a relational store, a message bus, and +an image registry. Together with +[ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md), which removed the +object store, the list is settled and every member is there for the same reason — the control +plane needs it and cannot ask itself for it. + +**A mesh that wants no identity provider runs none.** That is now expressible, and was not while +it sat in the substrate as a maybe. + +## Consequences + +**The last open question about substrate membership is closed.** Both halves of ADR 0006's test +now have an answer for every candidate, and the answer for identity is *the control plane does not +need it*. + +**It does not settle how a person signs in to a mesh surface**, and that is deliberately left +open. What is settled is that whatever answers it is not part of what must exist before the mesh +does — so it can be decided late, changed, or replaced, which is precisely what being substrate +would have prevented. + +**It becomes a real test of the module graph.** An identity provider is a module that *other +modules require* — the object store already consumes it — so it exercises the provider chain more +seriously than anything ported so far, where the provider was written alongside its consumer. + +**Ordering follows from it rather than from preference.** Anything requiring identity has to move +after it, which is a dependency the graph can state rather than something a person has to +remember. + +## References + +- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the conditional this closes +- [ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) — the other member + removed, and the test applied properly +- [ADR 0004](0004-a-node-and-how-it-joins.md) — how a node proves itself, which needs none of this diff --git a/02-DECISIONS/0031-the-lab-provides-the-underlay.md b/02-DECISIONS/0031-the-lab-provides-the-underlay.md deleted file mode 100644 index 4b69888..0000000 --- a/02-DECISIONS/0031-the-lab-provides-the-underlay.md +++ /dev/null @@ -1,87 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 31. The lab provides the underlay; the mesh builds the overlay - -## Context - -[Research 004](../01-RESEARCH/004-lab-network/analysis.md) worked out the topology a lab has to -reproduce: a routable segment using documentation addresses, a household segment behind NAT, a -router that forwards exactly one port so a *published-but-NATed* node is real, and a machine -that can attach to either segment or detach entirely. - -It also records what makes that topology **mean** something, and this is where a boundary has -to be drawn. Hub election is by convention rather than by flag — the hub is the node whose -profile is server and whose overlay address begins `10.10.0.1`. Direct peering depends on two -nodes sharing a site. Names resolve from mesh configuration on each node. - -Those are all facts the *mesh* establishes. The question is whether a scenario declares them. - -It is tempting to say yes, because a scenario that hands you a working overlay is a scenario -you can start testing against immediately. - -## Considered options - -1. **The lab configures the overlay too** — assign the overlay addresses, elect the hub, write - the peer configuration, seed the names. Rejected, and the reason is the whole point of the - lab: **a lab that builds the overlay certifies its own work.** If the mesh's peering logic - is broken, a scenario that pre-built the peering still comes up green. The most valuable - thing the lab can test is precisely the part this would replace. -2. **The lab provides nothing but bare machines** — no addressing, no segments, no NAT. Also - rejected. Then the scenario cannot reproduce *published but behind NAT*, which research 004 - identifies as the case that only exists in production today, and the lab loses its reason to - use virtual machines at all. -3. **The lab provides the underlay; the mesh builds the overlay.** Chosen. - -## Decision - -**A scenario declares the underlay** — the facts a machine would have before any of our -software touched it: - -- which segments exist, and their address ranges -- which machine sits on which segment, at which address -- what NAT sits between them, and which ports are forwarded through it -- which machines are detached, and can be attached or detached during a run - -**A scenario declares nothing about the overlay** — no overlay addresses, no hub, no peering, -no names, no certificates. Those are the mesh's job, and a scenario that supplied them would be -testing itself. - -The rule stated in one line: **a scenario provides what a hosting provider and a home router -would provide, and nothing our software is responsible for.** - -## Consequences - -- **The overlay becomes a thing under test rather than a fixture.** Whether peers form, - whether the hub is elected, whether a NATed node's endpoint is learned — all of it is - observed rather than arranged. That is the class of fault research 004 says is discoverable - only in production today. -- The lab stays small, and stays honest. It needs to know about virtualisation, bridges, - addresses and NAT. It never needs to know what a mesh node is. -- A scenario cannot assert "the overlay came up" as a precondition, because it is an outcome. - A bootstrap scenario that wants a working overlay has to wait for one and check, which is - the correct shape. -- **The address ranges are load-bearing, not cosmetic.** The routable segment uses RFC 5737 - documentation space specifically because the mesh's own code decides *public versus private* - by matching the address — a private range there makes the hub test as unreachable, and the - mesh silently never forms. Research 004 calls this the single most important fact in the - document, and the declaration format has to make getting it wrong hard. -- The router is a machine the lab materialises without being asked, because NAT requires - somewhere to run. That is an implicit machine in an otherwise explicit declaration, and it is - worth knowing about rather than discovering. -- **Host capability profiles are detected, not declared** — a consequence of - [research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md), and consistent here: a - scenario does not say what a machine is allowed to do, it provides a machine. Which leaves an - open question: a lab machine is always privileged, so the `user` and `edge` profiles have no - scenario that exercises them yet. - -## References - -- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology, the documentation - ranges, and the hub-election and peering conventions this deliberately does not touch. -- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the two scenario classes this - declaration has to serve without forking. diff --git a/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md b/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md deleted file mode 100644 index 09dac33..0000000 --- a/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md +++ /dev/null @@ -1,71 +0,0 @@ ---- -status: accepted -date: 2026-08-23 -deciders: jochen -reconstructed: false ---- - -# 32. A scenario is an isolated address space, and the lab never reaches into it over IP - -## Context - -A scenario declares literal addresses — -[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has -to be, because reproducing *published but behind NAT* means saying which address the world sees. - -That raises a question the declaration left open: **two scenarios at once.** Several agents -working means several scenarios, and the lab design already calls that a requirement. But two -scenarios built from the same declaration want the same addresses, and there are only three -documentation ranges in existence. - -## Considered options - -1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals. - Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a - specific topology no longer reproduces it; it breaks the RFC-range validation, since - allocated addresses would have to come from somewhere real; and the numbers a person reads - in the file stop being the numbers they will see in a capture. -2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate - an agent has to queue for is a gate that gets bypassed. -3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen. - -## Decision - -**A scenario is a closed address space.** Every segment materialises as its own isolated link, -belonging to one scenario instance. Two scenarios raised from the same declaration hold the same -addresses and never meet, because nothing joins their links. - -The declaration therefore keeps its literal addresses, and they mean exactly what they say. - -**The consequence that constrains everything else: the lab never reaches into a scenario over -IP.** It talks to a machine through the virtualisation layer's own channel — the same way one -executes a command in a container without the container being routable. - -That is not a preference. If the lab reached machines by address, the workstation running it -would need a route into each scenario, and two scenarios carrying the same prefix would give it -two routes to the same destination. Concurrency would be impossible, and it would fail in the -worst available way: not with an error, but by one scenario's traffic arriving in another. - -## Consequences - -- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit - beyond the machine's capacity. -- The three documentation ranges stop being a scarce resource. Every scenario may use all of - them, because no two scenarios share a link. -- **The lab cannot use IP to check anything**, which is more of a constraint than it first - appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by - executing on a machine, rather than probed from outside. That is the honest way to ask it - anyway — reachability from the workstation is not the question. -- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing - outside holds a reference into it. -- The lab needs a scenario **instance** identity distinct from the scenario name in the - declaration: the declaration is a kind, and several instances of one kind may exist. -- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a - developer wants to reach — a web interface, a database — needs an explicit, deliberate - forward out of the scenario, which is a feature rather than a gap: nothing leaks by default. - -## References - -- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses - this preserves. -- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes. diff --git a/02-DECISIONS/0032-the-local-account-owns-the-mesh.md b/02-DECISIONS/0032-the-local-account-owns-the-mesh.md new file mode 100644 index 0000000..81a3ced --- /dev/null +++ b/02-DECISIONS/0032-the-local-account-owns-the-mesh.md @@ -0,0 +1,79 @@ +--- +topic: how we work +status: superseded +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0031-the-control-plane-authenticates-nobody.md +superseded-by: 02-DECISIONS/0034-the-local-account-owns-the-mesh.md +--- + +# 32. The local account owns the mesh; a surface delegates to a module + +## Context + +[ADR 0031](0031-the-control-plane-authenticates-nobody.md) settled that the control plane +authenticates nobody, and deliberately left one thing open: **how a person signing in to a mesh +surface is authenticated.** This answers it, and answers a question 0031 did not ask — *who owns +the mesh at all.* + +**There was no answer, and the absence was invisible** because every operation so far has been run +by the person sitting at the machine. Nothing had to say whether that was the design or the +circumstance. + +## Decision + +**The account that installed the host owns the mesh on that node.** Authority is a local login, +and there is nothing else to hold. + +**No mesh user model.** No accounts, no roles, no grants, nothing to administer. A person with a +shell on a node can do anything the mesh can do there, because that is already true and pretending +otherwise would be a boundary that does not exist. + +**This follows from what was already decided rather than adding to it.** +[ADR 0004](0004-a-node-and-how-it-joins.md) says there is no authorisation between nodes — every +node is the operator's own, so a message from one is a message from them, and *the mesh boundary +is therefore the security boundary*. A user model inside that boundary would guard nothing: anyone +who could be stopped by it could equally read the node's key off the disk. + +**The board is different, and the difference is the network.** A surface reachable by a browser +has to know who is asking, because the people reaching it are not, by construction, people with a +shell on the machine. **So the board delegates to an OAuth provider** — which is a module. + +## What this does not change + +**The identity provider is still not substrate** (ADR 0031). A *surface* delegating +authentication is not *the control plane* delegating it. The control plane runs, applies +declarations and reaches nodes with no identity provider in existence; only the board needs one, +and only to decide whose browser it is talking to. + +The test is unchanged and still answers no: *does the control plane need it in order to run?* + +## Consequences + +**The board depends on a module, and says so.** An ordinary edge in the graph, which means the +board cannot come up before the provider it authenticates against — stated as a dependency rather +than discovered as an outage. + +**Moving the identity provider takes the board with it.** During that module's own conversion the +board is unavailable, and that is acceptable: it is a surface, nothing depends on it, and a brief +interruption is the trade already accepted everywhere else. Nothing that keeps a service serving +goes through it. + +**Anyone with a shell on a node has full authority there.** Written down rather than left implied, +because it is the sentence that decides who gets an account on a machine. The protection is the +machine's own login, and the overlay that keeps the machine unreachable from outside +([ADR 0007](0007-connectivity.md)). + +**A node cannot be operated by somebody without a login on it.** Deliberate, and the cost of +having no user model: there is no way to give a person authority over one node without giving them +a shell there. If that is ever wanted, it is a new decision and not a gap in this one. + +## References + +- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — the control plane authenticates + nobody; this answers what it left open +- [ADR 0004](0004-a-node-and-how-it-joins.md) — no authorisation between nodes, and why the mesh + boundary is the security boundary +- [`03-DESIGN/01-to-be/11-a-board.md`](../03-DESIGN/01-to-be/11-a-board.md) — the surface this is + about diff --git a/02-DECISIONS/0033-a-router-is-scenery-not-a-node.md b/02-DECISIONS/0033-a-router-is-scenery-not-a-node.md deleted file mode 100644 index a98bb85..0000000 --- a/02-DECISIONS/0033-a-router-is-scenery-not-a-node.md +++ /dev/null @@ -1,81 +0,0 @@ ---- -status: accepted -date: 2026-08-24 -deciders: jochen -reconstructed: false -extends: 0016-a-lab-node-is-a-virtual-machine.md ---- - -# 33. A router is scenery, not a node — so it is a container - -## Context - -[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) settles that **a lab node is a virtual -machine**, and its reasoning is fidelity: a node boots a stock image and runs the real install, -so it has to be a real machine or the thing under test is not the thing that ships. - -A scenario also needs routers. NAT, port forwarding, policy between segments and mapping -expiry are all things a router does, and until one is materialised a multi-segment scenario -raises isolated islands -([03-DESIGN/01-to-be/02-scenario-declaration.md](../03-DESIGN/01-to-be/02-scenario-declaration.md)). -The declaration already implies them: a gateway is *the one implicit machine in an otherwise -explicit declaration*. - -The question is whether ADR 0016 binds those too. - -## Considered options - -1. **A router is a node, so it is a virtual machine.** Consistent, and pays for a consistency - nobody needs. A router boots in roughly ten seconds against a container's one; a - six-segment scenario wanting three routers spends thirty seconds per raise on scenery. -2. **The hypervisor provides NAT** — bridges with translation switched on, and its own - forwarding primitives. Rejected on a stronger ground than speed: it makes the *lab* provide - what the declaration is supposed to own, and it cannot express a mapping that expires, a - gateway that refuses to forward, or policy between siblings. The model would shrink to fit - the tool. -3. **A router is scenery, and scenery is a container.** Chosen. - -## Decision - -**ADR 0016 binds nodes. A router is not a node.** - -Nothing under test runs on a router. It is not a participant, it holds no identity, the mesh -never installs anything on it, and no assertion is ever made about its internals. It exists so -that packets between machines behave the way they behave in the world — which is the definition -of scenery. - -So a router is a **system container**, and the fidelity argument does not reach it: what a -router must reproduce is kernel behaviour — translation, connection tracking, filtering, -forwarding — and a container has the same kernel. - -**Verified before deciding, not assumed.** In a plain unprivileged container: - -| Needed for | Works | -|---|---| -| routing at all | `net.ipv4.ip_forward`, `net.ipv6.conf.all.forwarding` | -| `nat:` | nftables masquerade, rules accepted and listed back | -| `mapping_ttl:` | `nf_conntrack_udp_timeout`, `nf_conntrack_tcp_timeout_established` | - -No privileged mode, no nesting, no capability grants. - -## Consequences - -- A raise stops paying a boot per router. Scenery costs about a second where a node costs ten, - and a scenario's cost tracks the machines actually under test. -- **The distinction is now load-bearing and has to stay legible.** *Node* means something under - test; *scenery* means something that makes the test real. If anything is ever installed on a - router by the mesh, it has become a node and this decision no longer covers it. -- Routers and nodes are different kinds of thing in the lab's own model, which is a small extra - concept — justified by it being true, rather than by the saving. -- A container shares the host kernel, so a scenario cannot reproduce a router running a - *different* kernel from the workstation. Nothing currently wants that; if something does, that - router becomes a virtual machine and this record needs revisiting rather than bending. -- The gateway stays implicit in the declaration. A scenario declares `gateway:` on a segment and - never names the machine that serves it — which is right, because it is not a machine the - scenario has anything to say about. - -## References - -- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — what a lab *node* is, unchanged. -- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology needing a router, - and why *published but behind NAT* only exists in production today. diff --git a/02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md b/02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md new file mode 100644 index 0000000..5de0e07 --- /dev/null +++ b/02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md @@ -0,0 +1,91 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md +--- + +# 33. The substrate is a store and a broker + +## Context + +Third correction to one table in one day, all found the same way: by asking whether **both** halves +of the substrate test were actually answered for a given member, or only the second. + +The test ([ADR 0006](0006-the-substrate-and-the-control-plane.md)) is *what the control plane needs +in order to run, and cannot ask itself for, because it is not running yet.* ADR 0006 admits the +image registry on this line: + +| role | product | | +|---|---|---| +| image registry | **an OCI registry** | it cannot grant itself a repository | + +**That is the second half again.** It is true that a control plane cannot grant itself a +repository. Nothing establishes that it needs one *in order to run*. + +**Counted rather than argued.** `substrate-first-node.lock` — the only bundle there is, and what a +first node actually becomes — raises twelve resources, and no registry is among them: + +``` +container runtime · the store · one database per context · the schemas +· the broker's certificate · the broker · the control plane +``` + +The registry arrives afterwards, as an ordinary module the mesh assigns. That is what the lab +asserts, in those words: *the mesh runs its own artifact store.* + +**ADR 0006 half-said this already**, calling the registry *substrate by role and ordinary by +delivery, provisioned once there is a control plane to do it.* A member that is provisioned by the +thing it supposedly precedes is not a member; the phrase was carrying a contradiction rather than +resolving one. + +**The registry is a closer call than the object store, and the difference is worth keeping.** The +control plane never touches an object store at all — no client, no bucket, ever +([ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md)). It genuinely +*uses* the registry: the builder pushes to it, hosts pull from it, and nothing reaches a machine +without it. **So the registry is a real dependency of the mesh operating, and not of the control +plane starting** — and it is the second that the word substrate means. + +## Decision + +**The substrate is two things: a relational store and a message bus.** Both are in the bundle, +both must exist before the control plane's first instruction, and neither can be asked for. + +**The registry is an ordinary module.** The mesh cannot deliver anything without one, and it +installs one the way it installs everything else. The first node's chicken-and-egg is already +solved and needs nothing from this list: it fetches upstream images directly, then runs a registry +of the mesh's own. + +**The test is applied to both columns, every time.** *Cannot grant itself one* is true of almost +any service and settles nothing on its own. It is what admitted the object store, and then the +registry, and both were removed by asking the other question. + +## Consequences + +**The substrate is now exactly what the bundle raises**, which is the strongest form this list can +take: it can be checked by counting rather than by reading an argument. A member that is not in +the bundle is not substrate, and the two statements cannot drift apart. + +**A mesh that builds nothing still needs a registry** — to receive anything at all — but it needs +it as a module, on its own schedule, replaceable. That was already true and was obscured by the +list. + +**The word may now be doing too little work.** "Substrate" for *a database and a broker* is a term +of art for two things everybody can name. Renaming is not taken here and is worth considering +separately; what this record fixes is the membership, not the vocabulary. + +**Three removals from one table in one day is itself the finding.** Each member was admitted on the +half of the test that is easy to answer, and the design read plausibly throughout. The rule that +comes out of it is not about substrates: **a test with two conditions is a test only when both are +asked.** + +## References + +- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the definition, and the table this + corrects a second row of +- [ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) — the object + store, removed for the same reason +- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — identity, which was conditional and + is now a module diff --git a/02-DECISIONS/0034-the-local-account-owns-the-mesh.md b/02-DECISIONS/0034-the-local-account-owns-the-mesh.md new file mode 100644 index 0000000..5c14d29 --- /dev/null +++ b/02-DECISIONS/0034-the-local-account-owns-the-mesh.md @@ -0,0 +1,81 @@ +--- +topic: how we work +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +supersedes: 02-DECISIONS/0032-the-local-account-owns-the-mesh.md +--- + +# 34. The local account owns the mesh, and a web application's login is not that + +*Supersedes [ADR 0032](0032-the-local-account-owns-the-mesh.md), which decided the right thing and +described it wrongly. The decision below is unchanged; what it said about the board was an +invention.* + +## Context + +ADR 0032 answered *who owns the mesh* — the account that installed the host — and then framed the +board as **a surface that delegates authentication**, a category it made up for the occasion. It +does not need one. + +**The board is a web application.** It has a login, provided by the identity module, in the way +every web application has a login. That is a fact about an application, not a property of the +mesh, and giving it a name in the mesh's vocabulary implied a relationship that is not there. + +The cost of the invented category was not cosmetic. It made the identity module look like part of +the mesh's own authority — something the mesh *depends on* to know who anybody is — when the truth +is that the mesh knows nothing about people at all, and one of the applications running on it has +a login. + +## Decision + +**The account that installed the host owns the mesh on that node.** Authority is a local login. +There is nothing else to hold, no user model, no roles, and nothing to administer. + +**This follows from what was already decided.** +[ADR 0004](0004-a-node-and-how-it-joins.md) says there is no authorisation between nodes — every +node is the operator's own, so a message from one is a message from them, and *the mesh boundary +is the security boundary.* A user model inside that boundary would guard nothing: anyone it could +stop could read the node's key off the disk. + +**A web application's login is its own business.** The board authenticates its users through the +identity module. So might anything else the mesh runs. **None of that is mesh authority**, and the +mesh does not learn who anybody is from it. + +## The line this draws, which is the reason to write it down + +**Signing in to an application must not, on its own, become authority over the mesh.** + +Today it cannot: the board reads and does not act +([`11-a-board.md`](../03-DESIGN/01-to-be/11-a-board.md) — *not the way to change things*). Looking +at a page tells you what is true and changes nothing. + +**The moment the board can assign a module, whoever it lets in has mesh authority** — and it would +arrive as a feature rather than as a decision. That is the failure this record exists to make +visible, because it is the kind that is only obvious afterwards. + +So: **a surface that can change the mesh is a change to who owns the mesh**, and is taken as one. +Not forbidden — wanting to manage nodes from a browser is reasonable — but not something that +turns up in a pull request titled *add assign button*. + +## Consequences + +**The identity module is not special.** Not substrate ([ADR 0031](0031-the-control-plane-authenticates-nobody.md)), +not part of the mesh's authority, and nothing about the mesh stops working when it is down. Some +applications cannot be logged into, which is what it means for an application's login provider to +be unavailable. + +**Anyone with a shell on a node has full authority there.** Unchanged from ADR 0032, and still the +sentence that decides who gets an account on a machine. The protection is the machine's own login +and the overlay that keeps it unreachable from outside ([ADR 0007](0007-connectivity.md)). + +**A node cannot be operated by somebody without a login on it.** The cost of having no user model. +If that is ever wanted, the paragraph above says what it costs. + +## References + +- [ADR 0032](0032-the-local-account-owns-the-mesh.md) — superseded; same decision, invented category +- [ADR 0004](0004-a-node-and-how-it-joins.md) — the mesh boundary is the security boundary +- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — the control plane authenticates + nobody diff --git a/02-DECISIONS/0035-one-implementation-several-surfaces.md b/02-DECISIONS/0035-one-implementation-several-surfaces.md new file mode 100644 index 0000000..722c076 --- /dev/null +++ b/02-DECISIONS/0035-one-implementation-several-surfaces.md @@ -0,0 +1,115 @@ +--- +topic: what runs on it +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0034-the-local-account-owns-the-mesh.md +--- + +# 35. One implementation, several surfaces, and what that costs + +## Context + +The mesh is operated from a command line today. It needs to be operable from a browser and from a +model's tools as well, and the three must not be three different systems. + +**The pattern is already in the code and unnamed.** `board` serves HTTP by calling the same +functions the CLI calls; it holds nothing and decides nothing. What follows makes that the rule +rather than a property of one command. + +**The board is a presentation layer over the control plane.** Not an application beside it holding +a database credential — the thing that shows what the control plane knows, and asks it to do what +a person asked for. + +## Decision + +**The logic lives once, in the context that owns it. A surface is an adapter with no decisions in +it.** + +| surface | for | +|---|---| +| **command line** | a person on a machine, and the recovery path below | +| **HTTP** | the board, and anything else that speaks to the mesh over a network | +| **model tools** | an agent asking the mesh to do something | + +**Every surface refuses identically, because the refusal is not in the surface.** An assignment +that cannot be satisfied is refused by the same resolution whichever way it arrived. The moment a +surface can accept something another would reject, the mesh has two answers to one question and +people learn which to trust. + +**Reading and doing are both exposed.** The HTTP surface is not read-only: managing the mesh from +a browser is the point. This takes the decision +[ADR 0034](0034-the-local-account-owns-the-mesh.md) said had to be taken deliberately — +**a browser login now carries authority over the mesh** — and takes it knowingly rather than +letting it arrive with a feature. + +**The networked surfaces authenticate through an OAuth2 identity provider.** Named by protocol +rather than by product, like every other dependency the mesh takes — AMQP for the bus, S3 for an +object store, OCI for the registry +([ADR 0006](0006-the-substrate-and-the-control-plane.md)). What fills the role today is a module +running Keycloak; what the control plane knows is that it validates a token against a provider +speaking OAuth2, and replacing that provider is a migration rather than a redesign. + +**The command line does not authenticate at all**: it is already behind the machine's own login, +which is what owns the mesh (ADR 0034). + +## What this is not: a kernel every module imports + +**The shared library is the failure this project was started over**, and the difference has to be +stated or it will be rebuilt. The old one is 155 files and 34,636 lines *containing code from +every context* — work-domain logic sitting in the kernel every module imports, each piece landing +there to avoid a cycle between two modules that both needed it. + +**Shared surfaces are not a shared library.** What is shared here is that three adapters call the +same functions. Those functions stay in the context that owns them — provisioning's logic in +provisioning, identity's in identity — and no module imports another's. A surface may call many +contexts; a context still may not reach into another's store +([ADR 0008](0008-a-context-owns-its-store.md)). + +The test, when something is about to be put "somewhere shared": *does this belong to a context, or +does it only belong to the surface?* If it belongs to a context it goes there, even if two +surfaces want it. + +## The loop this creates, and the way out + +**The control plane's networked surfaces will depend on a module the control plane assigns.** +An identity provider is an ordinary module ([ADR 0031](0031-the-control-plane-authenticates-nobody.md)). When +it is down, or being migrated, or misconfigured, the HTTP and tool surfaces cannot authenticate +anybody — including the person trying to fix it. + +**The command line is the way out, and it is why local ownership matters more rather than less.** +It authenticates through nothing, needs no network, and is available on the machine to the account +that owns the mesh. **A mesh must always be operable by somebody standing at it.** + +So the rule: **no capability exists only behind an authenticated surface.** Anything the board can +do, the command line can do. That is not a courtesy to CLI users; it is the recovery path, and a +capability that exists only over HTTP is one that disappears exactly when identity does. + +## Consequences + +**Identity is still not substrate**, and the test still answers no: the control plane runs, applies +declarations and reaches nodes with no identity provider in existence +([ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)). What is unavailable without it is two +surfaces, not the mesh. + +**Whoever the identity provider admits has authority over the mesh.** That is now a real perimeter +with real consequences, where before it guarded a page that only read. Who may log in, and to +which realm, becomes a decision about the mesh rather than about an application. + +**A surface must not grow an opinion.** The likely erosion is a validation added to the board +because it was quicker there — and then the CLI accepts something the board rejects, or worse the +reverse. Adapters hold no decisions. + +**Three surfaces over one implementation is a cost paid three times if it is not one +implementation.** The reason to write this down now is that the second surface is the cheapest +moment to get it right, and the third is where the drift usually starts. + +## References + +- [ADR 0034](0034-the-local-account-owns-the-mesh.md) — the local account owns the mesh, and the + line this record deliberately crosses +- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the shared library this must not + become +- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store, which a surface does + not change diff --git a/02-DECISIONS/0036-a-node-is-a-managed-machine.md b/02-DECISIONS/0036-a-node-is-a-managed-machine.md deleted file mode 100644 index 0cf74b3..0000000 --- a/02-DECISIONS/0036-a-node-is-a-managed-machine.md +++ /dev/null @@ -1,75 +0,0 @@ ---- -status: accepted -date: 2026-08-25 -deciders: jochen -reconstructed: false ---- - -# 36. A node is a managed machine, and disconnection is a situation - -## Context - -[Research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) left open: *"does an -unprivileged node earn a place in the inventory, or only a presence? Decides whether 'node' -means one thing or two."* - -The question came from requirement 6 — *Arch Linux only for now; ideally any device, including -phones, on lighter terms* — and from the observation that some machines cannot be fully -managed. A phone will not run the host. A laptop is absent for days. - -The question assumed the answer was a **class**: full nodes and lesser ones, with the -inventory recording the first and merely acknowledging the second. - -## Considered options - -1. **Two classes — nodes and presences.** An unprivileged device gets a lighter record and a - reduced contract. Rejected: it makes "node" mean two things, so every context that reasons - about nodes acquires a branch, and the branch is invisible in the type. The mesh already has - one instance of this shape and it is the one this repository keeps writing issues about — - a declared thing that is only sometimes honoured. -2. **One class, membership by capability.** Everything is a node; what it can do is a property. - Chosen. - -## Decision - -**A node is a managed machine inside the mesh.** Not a device that is merely known about, not -an unprivileged something. If the mesh does not manage it, it is not a node — it is a client, a -peer, or a thing on the network, and those want their own names rather than a weakened version -of this one. - -**A disconnected node is still a node, in a different situation.** Reachability is state, not -class. A node that is switched off, roaming, or behind a connection that has dropped has not -become a lesser kind of thing; it has a last-known state and a pending set of declarations. - -The distinction the original question reached for is real, but it is **capability**, not kind — -what this machine can be asked to do — and that belongs in the host's profile, not in the -definition of a node. - -## Consequences - -- **The inventory has one shape.** No branch, no second record type, no context that must ask - which kind it is holding. -- **Local state is structural, not a convenience.** If disconnection is an ordinary situation - rather than an exception, the host's store is authoritative while disconnected by design — - it is what makes the situation ordinary. This promotes `store/` from a component to a - requirement. -- **Absence is not failure.** A node that has not been seen is in a state, and the mesh must be - able to say which. Anything that treats unreachable as broken will be wrong most of the time - about a laptop. -- **Devices that cannot be managed do not become nodes by being lenient about the word.** A - phone that cannot run the host is not a node under this record. Whether the mesh should reach - such devices at all, and as what, is not decided here and needs its own record if it is - wanted. -- **The reduced-contract idea is not lost, it is relocated.** What a given node can be asked to - do is its profile — the host's capability detection — and varies per machine without varying - what a node is. - -## References - -- [Research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) — the open question, and - requirement 6 that raised it. -- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — *nodes host*; this says what a - node is. -- [Issue 007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) — - capability as something detected rather than assumed, which is where the reduced contract - now lives. diff --git a/02-DECISIONS/0036-bootstrap-ends-at-a-usable-mesh.md b/02-DECISIONS/0036-bootstrap-ends-at-a-usable-mesh.md new file mode 100644 index 0000000..87d330d --- /dev/null +++ b/02-DECISIONS/0036-bootstrap-ends-at-a-usable-mesh.md @@ -0,0 +1,99 @@ +--- +topic: the tiers +status: accepted +date: 2026-08-31 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0035-one-implementation-several-surfaces.md +--- + +# 36. Bootstrap ends at a usable mesh, and the first credential comes from a person + +## Context + +Bootstrap currently ends when the control plane starts +([`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md)). That is a mesh that runs and +cannot yet be used by anybody who is not standing at the machine: the networked surfaces need an +OAuth2 identity provider ([ADR 0035](0035-one-implementation-several-surfaces.md)), the provider is +a module, and no module has been assigned. + +**So bootstrap should go further** — through the identity provider and the first login — and stop +at a mesh somebody can actually use. + +**One thing in the way, and it is not incidental.** The mesh has never held a readable secret. The +sealing code says what it does and why: + +> Make generates a secret and seals it to both ends, **keeping no readable copy.** + +An initial administrator's credential is the first value a **person must read**. Everything else +the mesh generates is something no human ever sees, and everything a human provides is something +the mesh immediately stops being able to read. + +## Considered Options + +1. **The mesh generates it and prints it once**, to the terminal of whoever ran the bootstrap. + Convenient, and needs no prompt. **Rejected.** It would give the control plane a plaintext + secret for the first time — briefly, and only to one terminal, but the capability would then + exist. *An exception made for one case does not stay one*: the next credential that is awkward + to supply gets printed too, and the property that a copy of the mesh's database is a copy of + nothing stops being checkable by reading the code. + +2. **No password: a one-time link that lets the operator set their own.** The nicest to use. + **Rejected for now** — it needs a mechanism that does not exist, and the thing it improves is + one prompt, once, on a new mesh. + +3. **The operator supplies it.** **Adopted.** + +## Decision + +**Bootstrap runs to a usable mesh**: the substrate, the control plane, the identity provider as an +ordinary module, its realm and client provisioned, an administrator able to log in, and the +networked surfaces available. + +**The administrator's credential is supplied by the person doing the bootstrap**, on standard +input and not echoed — the path that already exists for a model-access key. The mesh seals it and +cannot read it afterwards. + +**What is created is an account in the identity provider, not a user of the mesh.** The mesh still +has no user model and gains none here ([ADR 0034](0034-the-local-account-owns-the-mesh.md)). What +this produces is the first login for the applications that have one. + +**The provisioning is ordinary.** A realm, a client and a first account are what an identity +module's provisioner makes from what the mesh granted it — the same shape as a database and a +bucket, which are built and proven. + +**The surfaces arrive when their dependency does.** The command API is not started with the +control plane and then broken until identity exists; it becomes available once it can authenticate, +the way anything else waits for a provider. + +## Consequences + +**The mesh still never holds a readable secret**, and that sentence needs no exception clause. +That is the whole reason for the prompt. + +**An unattended bootstrap is still possible, and the value still comes from outside.** Automation +supplying the credential is the operator supplying it. What is refused is the *mesh inventing* +one — so an unattended bootstrap with no credential provided produces a mesh with no +administrator, which is correct rather than broken. + +**Bootstrap gains an interactive step**, and it is the only one. Worth stating because a bootstrap +that cannot run without a person is a real constraint on how a node is stood up, and this is +deliberate rather than an oversight. + +**The identity provider is still not substrate.** It is assigned by the control plane, so it comes +after it, and a thing that comes after cannot be a thing that must exist before +([ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)). Bootstrap running through it does not +move it: bootstrap is a sequence, the substrate is a dependency. + +**And the recovery path is unchanged.** When the identity provider is broken later — which is the +failure that matters, not the one at first start — the command line still works, because it +authenticates through nothing (ADR 0035). + +## References + +- [ADR 0035](0035-one-implementation-several-surfaces.md) — the surfaces, and why the command line + must keep working +- [ADR 0034](0034-the-local-account-owns-the-mesh.md) — the local account owns the mesh; this adds + no user model +- [ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md) — what must exist before the control + plane, which this does not change diff --git a/02-DECISIONS/0037-the-host-applies-it-does-not-decide.md b/02-DECISIONS/0037-the-host-applies-it-does-not-decide.md deleted file mode 100644 index 301a318..0000000 --- a/02-DECISIONS/0037-the-host-applies-it-does-not-decide.md +++ /dev/null @@ -1,97 +0,0 @@ ---- -status: accepted -date: 2026-08-25 -deciders: jochen -reconstructed: false ---- - -# 37. The host applies; it does not decide - -## Context - -The skeleton absorbs overlay membership, packet filtering, package management, service -supervision, the container runtime and filesystem management into tier 0, and -[research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) called this *"the -skeleton's biggest unproven claim. A binary whose whole argument is that it has no dependencies -now carries six concerns."* - -That claim has now been measured against the monorepo's `main`: -[`host-size.md`](../01-RESEARCH/006-mesh-from-scratch/host-size.md). - -The measurement says the question asked about the wrong axis. - -## Considered options - -1. **Absorb the six concerns as they are.** What the skeleton literally proposes. Rejected on - evidence: two of the ten modules implementing them open a direct connection to the control - plane's database and compute their own configuration. Absorbing those unchanged puts a - Postgres client and knowledge of the mesh schema inside tier 0 — an upward dependency, and - the tier rule is the whole of the bootstrap argument. -2. **Leave them as modules.** Keeps the tier rule trivially, and keeps the fault that prompted - the skeleton: four modules constituting *how a node is reachable* with no relationship the - mesh can see, so one intent is expressed four times - ([research 005](../01-RESEARCH/005-domain-grouping/analysis.md) finding 4 measures this and - finds it is the only place in the catalogue where the shape genuinely occurs). -3. **Split each concern: decide centrally, apply locally.** Chosen. - -## Decision - -**The host has one concern: apply declared state on this machine.** The six are not six -concerns it carries; they are instances of the one. - -Each divides: - -- **Deciding** — what this node's overlay, names, exposure, filtering, packages and services - *should be*. This needs every other node, and belongs to the control plane. -- **Applying** — putting that on the machine. This needs root and locality, and belongs to the - host. - -**The host never queries the mesh database.** A host that reads the control plane's schema is -tier 0 depending on tier 2, and the tiers stop being a bootstrap answer the moment that is -permitted once. - -## Why the evidence supports it - -**Size was the wrong worry.** The ten modules total 2 755 lines. The machinery that already -applies state on a node — `meshware`, `env-sync`, `config-sync` — is 3 059. Everything being -absorbed is smaller than what already exists to apply it. The host is not a new large thing; it -already exists, spread across three core modules and unnamed. - -**Eight of the ten are already pure appliers.** They receive derived state and put it on the -machine. Absorbing them moves code that has no dependency to move. - -**The split has already been happening, unnamed.** `dnsmasq-app` needs the same mesh-wide data -as `wireguard` and does not query for it. Its own comments record why: the values were -*"duplicated by hand on all four nodes"* until someone derived them centrally, after a rename -meant editing four override rows nobody knew about. That is this decision, reached once by -fixing a bug. - -## Consequences - -- **Two modules must be split before they can be absorbed**, and they are the two hardest. - `wireguard` needs every node's key, address, site and endpoint reachability; `traefik` needs - certificates and every node's exposed names. The measurement says the design is right; it - does not say the migration is cheap, and this record does not claim it is. -- **The overlay and firewall modules stop existing** as the skeleton says — but the reason is - now sharper than "they are host concerns". The host holds membership and applies filtering; - the control plane decides policy; swappable backends stay modules. -- **Six vocabularies remain.** Zero dependencies, but the host must still know what a WireGuard - peer, an nftables rule, a package, a unit, a container and a dataset *are*. That surface is - the residue of the original worry and is not measured by anything here. -- **A dependency-direction lint is now load-bearing**, not a nicety. This record is a rule - about direction, and per this repository's own standard a rule states how it is checked: an - upward import fails the build. A tier rule enforced by intention is the same as no tier rule. -- **What the host carries versus what it finds is still open.** - [Issue 007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) — the - host manages `wg`, `nft`, `pacman`, `docker`; it does not contain them, and *installed* is - not the same as *usable*. - -## References - -- [`host-size.md`](../01-RESEARCH/006-mesh-from-scratch/host-size.md) — the measurement. -- [Research 005](../01-RESEARCH/005-domain-grouping/analysis.md) — reachability as the only - measured co-change cluster in the catalogue. -- [ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) — what the control plane decides - from. -- [ADR 0030](0030-the-repository-structure.md) — `mesh-host` as tier 0. -- [ADR 0008](0008-a-failed-step-fails-the-job.md) — the standard the direction lint is held to. diff --git a/02-DECISIONS/0037-where-a-module-lives.md b/02-DECISIONS/0037-where-a-module-lives.md new file mode 100644 index 0000000..a697d12 --- /dev/null +++ b/02-DECISIONS/0037-where-a-module-lives.md @@ -0,0 +1,103 @@ +--- +topic: building it +status: proposed +date: 2026-09-01 +deciders: jochen +reconstructed: false +rests-on: 02-DECISIONS/0009-modules-and-the-graph.md +--- + +# 37. Where a module lives + +## The question + +The mesh's own module descriptions currently sit in `examples/` inside the control plane, beside +the small programs that hand out logins. That was fine while there were three of them. It is +wrong now, and the name is doing active harm: everything in `examples/` reads as a sketch, and one +of them shipped naming a container image that nothing in the repository builds. A directory called +*the catalogue* would have made *does this actually work* the obvious question to ask of it. + +So: **one repository holding the modules we ship?** And if so, where does everything that is not +ours go? + +## What a module actually is, counted + +The system being replaced has **126 modules** on its main branch. The shape of them is the whole +argument, so it is measured rather than assumed: + +| | count | what it is | +|---|---|---| +| **the module is software** | 47 | its own source tree lives inside the module — a daemon, a service, a library | +| **helper scripts only** | 44 | no application of its own; scripts it runs at install time or offers to an agent | +| **a description and nothing else** | 35 | a package to install and some files to write | + +**Two thirds of modules contain code.** The largest is a shared library of 182 source files. A +speech-capture module carries a complete daemon — audio capture, mixing, transcription, a model +runner. Treating a module as *a description of something else* is true of barely a quarter of them. + +That kills the simplest answer. A catalogue cannot be "a folder of manifests" when most modules +are programs. + +## The four kinds, which want different homes + +**1. What the mesh is made of.** The control plane, the host, the shared library, the board. +These are not modules that happen to be ours; they are the mesh, expressed as modules so it can +install itself. They belong in the repositories that build them, which already exist. + +**2. Something the world made, that we describe.** A forge, a mail system, an identity provider, +a media server. Nobody upstream ships a description; somebody has to write one, and it is the same +description for everybody who runs it. **This is what a catalogue is for.** It is also where the +small programs that create accounts belong, because such a program is part of describing that +service, not part of the mesh. + +**3. Something we wrote, that runs somewhere.** An application, a site, a side project. The +description belongs **with the code, at the root of its own repository**, because the two change in +the same commit. A repository that gains an environment variable and a description that gains it +elsewhere will drift, and there is no mechanism that could stop it. This is already how it works +and it should stay that way. + +**4. A package and some files.** A tool, a font, a shell. Thirty-five of these, and each is a few +lines. The catalogue. + +## The proposal + +**A `mesh-catalog` repository** holding kinds 2 and 4: descriptions of software we did not write, +and the programs that provision it. Not kind 1, which is the mesh itself. Not kind 3, which lives +with its own code. + +**The mesh's list of modules is not this repository.** It is a table in the control plane, filled +by adding a description to a running mesh. The catalogue is a *source* to add from — one of +several, and the mesh already records which: every module carries where it came from, the branch +followed there, and the commit its description was read at. **Nothing needs inventing to support +modules from anywhere**; a repository of our own is simply the source we curate. + +**A description is checked by the tool, not by a test that imports the tool.** Today a test in the +control plane parses the example manifests by reaching into the control plane's internals, and +another reads the control plane's own build file to check every image a module names can be built. +Two jobs tangled. A `module check` command on the control plane's binary would let the catalogue +hold data validated from outside, and would give the same check to somebody describing their own +application in their own repository — which is the case that matters most and currently has no +check at all. + +## What this costs, and the argument against + +**It is early.** Ten modules exist, four of them ours. Moving ten files is a morning; moving a +hundred is a week — but the hundred is not here yet, and splitting now adds a second repository to +release across before there is anything to release. + +The counter is that the tangle is already producing faults rather than merely threatening to. A +manifest naming an unbuildable image, and a test reading a build file two directories up, are both +symptoms of one repository doing two jobs. And the moment the first module is adopted on a real +machine, the descriptions stop being examples and become the thing deployments come from. **That +is the moment this becomes urgent, and it is close.** + +## What it does not settle + +**Where a provisioning program's image is published**, and how a description pins it. A description +names an image by digest; the image is built from the catalogue; the catalogue must therefore both +produce an image and refer to it, which is the same knot the bootstrap has and solves by writing +the digest down after building. + +**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and +write these files* may be better as one module with settings than as thirty-five modules. Left +open deliberately; it is a question about the shape of the catalogue, not about whether to have one. diff --git a/02-DECISIONS/0038-a-node-joins-by-linking-first.md b/02-DECISIONS/0038-a-node-joins-by-linking-first.md deleted file mode 100644 index e0c0f22..0000000 --- a/02-DECISIONS/0038-a-node-joins-by-linking-first.md +++ /dev/null @@ -1,106 +0,0 @@ ---- -status: accepted -date: 2026-08-25 -deciders: jochen -reconstructed: false -extends: 0037-the-host-applies-it-does-not-decide.md ---- - -# 38. A node joins by linking first, and the mesh finishes the job - -## Context - -[ADR 0037](0037-the-host-applies-it-does-not-decide.md) settles that the host applies and the -control plane decides. That leaves the case where there is no control plane to decide: the -first node, which must raise a mesh from nothing, and the second, which must join one. - -Raised by the operator: *"shouldn't the host have two modes — one for the initial node, setting -up the mesh, so we know the full state; then when adopting a second node, we enter the mesh -early and let our first node take over the mesh-related work? The host should only set up the -bare minimum for the other nodes in the mesh to complete adoption."* - -The instinct is right and it is the resolution of the gap ADR 0037 leaves open. The framing -needs one correction, and the correction comes from what the mesh already does. - -**Today there are three hand-run shell paths**: `install.d/adopt.sh` (183 lines), -a separate first-node bootstrap, and `install.d/rescue.sh`. The skeleton names the cause — -*"the first node is raised by a special script that exists only because of the circularity"* — -and Move 1 exists to remove it. Three paths that do nearly the same thing, maintained -separately, run by hand, outside anything that checks them. - -**That is the two-mode problem, already at its worst.** A decision that gives the host two -modes risks rebuilding `adopt.sh` and `bootstrap.sh` inside the binary, where they will drift -in exactly the same way and be harder to see. - -## Considered options - -1. **Two modes — genesis and join.** What was proposed. Rejected as a *structure* while adopted - as an *intent*: two modes is two code paths, the first is exercised once per mesh and the - second constantly, so the rarely-run one rots. The current three scripts are the evidence. -2. **One path, and the first node is special-cased inside it.** The conditional moves rather - than disappearing, and now it is scattered instead of named. -3. **One behaviour, two sources of declaration.** Chosen. - -## Decision - -**The host has one behaviour: apply the declaration it has.** What differs between the first -node and the fiftieth is not what the host *does* but **where the declaration comes from** — -and, exactly as in [ADR 0036](0036-a-node-is-a-managed-machine.md), that is a situation rather -than a class. - -| Situation | Declaration comes from | -|---|---| -| no mesh reachable | the pinned bundle the host carries (`substrate.lock`) | -| mesh reachable | the control plane, over the link | - -**The first node is not a different kind of node.** It is a node whose mesh is not up *yet*. It -applies the bundle it carries, the control plane comes up on top of it, and from that moment it -takes declarations like everything else. Its specialness is temporary and self-erasing, which -is the property `adopt.sh` and the bootstrap script do not have. - -**A joining node does the minimum to be reachable, and nothing else.** It establishes identity -and a route to the control plane — the `link` — and then stops deciding. Everything after that -arrives as declarations. - -**The minimum is deliberately small:** an identity, an address, and one peer to reach. A -joining node does **not** compute the overlay. It needs a single peer to reach the mesh; the -full peer set is derived centrally and pushed down, like everything else. - -## Why this resolves what 0037 left open - -ADR 0037 records that `wireguard` and `traefik` are the two modules that must be split before -they can be absorbed, and that they are the hardest because they need mesh-wide state. - -**A joining node never needs that state.** The hard part of the overlay — every node's key, -address, site and endpoint reachability — is only needed to compute the *whole* mesh, which is -the control plane's job. The node needs one peer. The rest arrives. - -So the migration ADR 0037 calls expensive is smaller than it looked, and this record is what -makes it smaller. - -## Consequences - -- **Adoption stops being a script.** The three hand-run paths collapse into the host: joining - is establishing a link, and rescue is a node whose local state is discarded so the mesh can - re-derive it. Whether rescue is fully covered by this is not decided here. -- **The bundle is a fallback, not a mode.** It is what a host applies when nothing better is - available, which also covers a node that has been disconnected for a long time — ADR 0036's - ordinary situation. -- **The rarely-run path is now the common one.** The first node exercises the same code every - other node exercises constantly. That is the whole reason for choosing this over two modes. -- **The link becomes the security boundary.** Everything a node applies arrives through it, so - what may be pushed, and how a joining node proves it is entitled to join, is its own - question — taken up by [ADR 0039](0039-the-link-is-the-security-boundary.md). -- **The bundle must be able to raise the substrate alone.** Whether one host can bring up the - four pinned services with no mesh present is Move 1 of the skeleton and remains unproven. - This record depends on it and does not establish it. - -## References - -- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — the split this completes. -- [ADR 0036](0036-a-node-is-a-managed-machine.md) — situation rather than class, applied here - to the first node. -- [Research 006, Move 1](../01-RESEARCH/006-mesh-from-scratch/skeleton.md) — the pinned bundle, - and the special script it exists to remove. -- [`00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md) - — how a node comes into being today. diff --git a/02-DECISIONS/0038-the-mesh-assigns-the-port.md b/02-DECISIONS/0038-the-mesh-assigns-the-port.md new file mode 100644 index 0000000..4b51302 --- /dev/null +++ b/02-DECISIONS/0038-the-mesh-assigns-the-port.md @@ -0,0 +1,82 @@ +--- +topic: what runs on it +status: proposed +date: 2026-09-01 +deciders: jochen +reconstructed: false +rests-on: 02-DECISIONS/0009-modules-and-the-graph.md +--- + +# 38. The mesh assigns the port, and a module does not care + +## The problem, as met + +A database module cannot start on a machine that runs the control plane. The mesh keeps its own +store there and holds 5432; the module publishes 5432. Nothing notices until a container runtime +three layers down says `port is already allocated` +([`028`](../04-ISSUES/028-two-things-want-one-port-and-nothing-says-so/00-report.md)). + +A module cannot fix this by choosing better, because **a module cannot know what else is on the +machine.** It is written once and assigned anywhere. Any number it picks is a guess about a +machine it has never seen, and two modules guessing the same number is not a mistake either of +them made. + +## The number is written three times, and nothing makes them agree + +Every module says its port in three places: + +| where | for | example | +|---|---|---| +| `listens` | the rule set that lets traffic in | `{port: 5432, from: mesh}` | +| `serves` | what a consumer must know to connect | `{port: 5432}` | +| a container's `ports` | what the runtime publishes | `"5432:5432"` | + +They agree today because one person wrote all three. Nothing checks it. A module whose `serves` +said 5432 and whose container published 5433 would resolve, compose, apply, and hand every +consumer a port that answers nothing. + +## The decision + +**The mesh assigns the machine-side port, and the module says only what it needs.** A module +declares that a container port must be reachable and what it is for. Which number the machine uses +is the mesh's to choose, because the mesh is the only thing that knows what else is there. + +**One source, and the other two are derived.** `serves` carries the assigned port so a consumer is +told where to connect without the module having written it down; the rule set is computed from the +same assignment. Three copies become one fact. + +**An assignment is made once and kept**, exactly as a credential is. A port that moved on every +push would restart both ends each time and would hand consumers a number that was true when it was +read. + +## Some ports cannot move, and that is a claim + +Mail is 25, submission is 587, IMAP over TLS is 993. A mail system on a strange port is not a mail +system. So a module may say a port is **fixed by the protocol** rather than assigned. + +**A fixed port is exactly a claim** — the thing the mesh already has for what is singular on a +machine: one seat, one display server, one artifact store. Two modules wanting 25 on one machine is +the same shape as two wanting the seat, and gets the same answer: the second is refused, by name, +when it is assigned rather than when it is applied. + +That is why this does not need a new mechanism so much as it needs the existing one pointed at +ports. + +## What follows + +- **A module becomes portable in a way it was not.** Two databases on one machine stop being a + collision and become two assignments. +- **The substrate has to be visible.** The mesh cannot assign around its own store while it has + never heard of it. What the bundle holds must be written down somewhere the assignment can read + — which the bundle does not say today. +- **A refusal can be useful.** *25 is held by the mail system on this machine* is a sentence a + person can act on. `port is already allocated` is not. +- **`serves` stops being written by hand**, which is a small vocabulary change with a large + consequence: what a consumer is told is now derived from what actually happened. + +## What this does not settle + +**Whether a module should publish to the machine at all.** Assignment makes publishing safe; it +does not make it necessary. Consumers could instead reach a provider on the module's own network by +name, with nothing published — which would make the question moot for anything inside the mesh, and +would still leave it for anything reached from outside. diff --git a/02-DECISIONS/0039-the-link-is-the-security-boundary.md b/02-DECISIONS/0039-the-link-is-the-security-boundary.md deleted file mode 100644 index 0d0e4de..0000000 --- a/02-DECISIONS/0039-the-link-is-the-security-boundary.md +++ /dev/null @@ -1,190 +0,0 @@ ---- -status: accepted -date: 2026-08-25 -deciders: jochen -reconstructed: false -extends: 0038-a-node-joins-by-linking-first.md ---- - -# 39. The link is the security boundary - -## Context - -[ADR 0038](0038-a-node-joins-by-linking-first.md) makes the link the one channel a node takes -declarations from, and names the gap it leaves: *"everything a node applies arrives through it, -so what may be pushed, and how a joining node proves it is entitled to join, is now a question -worth its own record."* - -This is that record. It is a design decision about a boundary that does not exist yet — but -what it replaces is measured, and that is the argument. - -Settled as: **a node owns no password. It owns an identity, and that identity is what it -presents to the broker.** - -### What adoption does today - -`install.d/adopt.sh` asks the operator to paste credentials in by hand: - -``` -The meshware module needs registry database and minio credentials. - REGISTRY_DB_PASSWORD= - REGISTRY_MINIO_PASSWORD= -``` - -plus an `NPM_TOKEN` for the private registry. These are not adoption-time credentials that are -then discarded: `wireguard` and `traefik` open a `pg` connection on every reconcile -([ADR 0037](0037-the-host-applies-it-does-not-decide.md)). - -**So every node permanently holds a credential to the control plane's database, and to the -object store.** They are the same credentials on every node. There is no rotation — -[`00-as-is/06`](../03-DESIGN/00-as-is/06-configuration-and-secrets.md) records that *"there is -no mechanism that rotates one and informs everything holding it. Where a rotation has been -done, it has been done by hand, and doing it wrong has taken services down."* - -Compromise of any node is therefore compromise of the mesh's database, and there is no -mechanism to recover from it. - -### The link is not new - -Written first as though the link were a thing to build. It is not. -[ADR 0001](0001-nodes-communicate-over-a-broker.md) already has it: *every node connects -outbound to a single broker; nothing ever connects to a node*, each node declaring an exchange -named for itself and consuming from its own queue -([`00-as-is/01`](../03-DESIGN/00-as-is/01-mesh-and-transport.md)). - -That is already outbound-only, already per-node addressed, and already the one channel -everything arrives through. **This record is not proposing a channel. It is proposing that the -channel carry per-node identity instead of one shared credential.** - -The same as-is records the fault, for the broker rather than the database: *"the broker is a -single point of failure and a single point of trust. Its credential is mesh-wide, so rotating -it is a mesh-wide operation, and doing it wrong has taken the broker down."* - -## Considered options - -1. **Keep shared credentials, scope them per node.** Least change: give each node its own - database role. Rejected — it makes the blast radius smaller without changing its shape, and - it keeps tier 0 speaking the control plane's schema, which ADR 0037 forbids for reasons that - are not about security at all. -2. **Accept the exposure as the cost of simplicity.** A shared credential is one thing to - understand and nothing to build, and the objection to replacing it is real: mutual - authentication fails opaquely, and a node that cannot link is harder to debug than a node - with a wrong password. Rejected on the ground that the simplicity is what makes it - unrotatable — the credential cannot be changed *because* everything holds the same one, so - the arrangement's convenience and its unfixability are the same property. -3. **Mutual authority on a node-initiated link, with the node holding nothing but its own - identity.** Chosen. - -## Decision - -**The link is the only way anything reaches a node**, and four properties make it a boundary -rather than a pipe. - -### It is outbound and node-initiated - -The node dials the control plane. Nothing dials a node. This is not only defensive — it is what -the topology already requires: most nodes sit behind a household connection with no forwarded -port ([research 004](../01-RESEARCH/004-lab-network/00-overview.md)), so an inbound control -channel would work for the hosted node and not for the rest, and the difference would be -invisible until it mattered. - -A node therefore has **no listening control surface at all**. - -### A node holds its own identity and nothing else - -No shared secret, no credential to anything it does not own. A node's identity authenticates it -to the control plane and grants access to nothing else. - -**Compromise of a node is compromise of that node.** That is the property today's arrangement -does not have, and it is the main reason for this record. - -### Authority is mutual - -The node proves it may join, and **the control plane proves it is the mesh**. One-way is not -enough here: the host applies whatever the link delivers, so a node that cannot tell the mesh -from something impersonating it will apply that something's declarations. Given ADR 0038, an -attacker who can answer a joining node's first call owns the machine. - -### What may be pushed is bounded by form, not by trust - -The control plane may push **declarations of known shape** and nothing else. It may not push a -command to run. The host's vocabulary is finite, versioned and auditable, and anything outside -it is refused rather than best-effort interpreted. - -**Stated honestly: this bounds form, not impact.** A compromised control plane can declare -harmful state — a malicious package, an open firewall — and the host will apply it faithfully, -because that is what it is for. What the property buys is that the blast radius is describable: -it is exactly what the declaration language can express, which can be reviewed. An arbitrary -command channel has no such bound. This is a real limit and not a defence-in-depth story. - -### Joining is a deliberate, bounded act - -A joining node presents a **one-time, short-lived enrolment token** issued by the mesh for that -purpose, and exchanges it for its own durable identity. The token grants exactly one thing: -the right to become a node. It is not a credential to any service, it does not persist after -exchange, and it expires whether used or not. - -This replaces hand-carried shared secrets with a thing that is useless once used and useless -after a while. - -## What this actually costs - -The objection to weigh is overhead, and it is smaller than it looks because most of it is -already running. - -| Property | Where it comes from | -|---|---| -| outbound, node-initiated | already true — ADR 0001 | -| per-node addressing | already true — per-node exchange and queue | -| per-node credential | a broker user per node; the broker already has users, virtual hosts and per-queue permissions | -| mutual authority | transport-level certificates on a connection that already exists | -| bounded by form | already true — three message shapes and only three | -| **enrolment** | **the one genuinely new mechanism** | - -And ADR 0037 subtracts rather than adds: under it a node holds **no** database credential at -all, so this record replaces three hand-carried shared secrets with one per-node identity that -grants only identity. - -**It must fail legibly.** A boundary that refuses a node without saying why is worse than the -credential it replaced, because a wrong password at least announces itself. A node that cannot -link must report which side rejected it and on what grounds, in terms someone can act on. This -is `how-we-build` §5 applied to a security mechanism: a refusal that proves only that something -went wrong is transport reported as effect. - -## Consequences - -- **ADR 0037 removes a standing exposure as a side effect.** Its rule — the host never queries - the mesh database — was chosen for tier discipline. It also removes the reason every node - holds the database password. Worth recording because the two arguments are independent and - both hold. -- **Rotation becomes possible and is still not designed.** Per-node identities can be revoked - individually, which is what makes rotation tractable at all. The mechanism — - what rotates, on what trigger, and how holders learn — is **not decided here** and remains - the open weakness `00-as-is/06` records. -- **The enrolment token has to come from somewhere.** Issuing it is a control-plane operation - and the first node has no control plane, so the first node's identity is self-issued and - becomes the root of trust when the mesh comes up. **That is a real asymmetry** — the one - place ADR 0038's "no special first node" does not fully hold — and it is named here rather - than hidden. -- **A declaration vocabulary is now a security artefact, not only a design one.** Every - addition widens what a compromised control plane can express. That is a reason to keep it - small and a reason for additions to be reviewed as such. -- **Offline nodes need identities that survive disconnection.** Per - [ADR 0036](0036-a-node-is-a-managed-machine.md) disconnection is ordinary, so an identity - that must be refreshed to remain valid would make a laptop fail for being a laptop. What - expires and what does not is **not decided here**. -- **This is a boundary that does not exist yet.** Nothing in the current mesh implements any of - it, and the migration from shared credentials to per-node identity touches every node and the - substrate. No estimate is offered. - -## References - -- [ADR 0038](0038-a-node-joins-by-linking-first.md) — the link, and the gap this fills. -- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — why the host stops holding database - credentials at all. -- [ADR 0036](0036-a-node-is-a-managed-machine.md) — disconnection as ordinary, which constrains - what may expire. -- [`00-as-is/06-configuration-and-secrets.md`](../03-DESIGN/00-as-is/06-configuration-and-secrets.md) - — secrets today, and the absence of rotation. -- [Research 004](../01-RESEARCH/004-lab-network/00-overview.md) — why most nodes cannot accept - an inbound connection. diff --git a/02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md b/02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md new file mode 100644 index 0000000..4199b09 --- /dev/null +++ b/02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md @@ -0,0 +1,106 @@ +--- +topic: building it +status: accepted +date: 2026-09-03 +deciders: jochen +reconstructed: false +--- + +# 39. What the SDK holds, and what it refuses + +_Reconciliation note (2026-09-05): supersedes the earlier "repository structure" decision, which the consolidation folded; no standalone record remains to point at, so body references to it now point at the nearest surviving record, [ADR 0015](0015-applications-live-in-their-own-repository.md)._ + +## Context + +The earlier "repository structure" decision (folded in consolidation; see the reconciliation note +above, and [ADR 0015](0015-applications-live-in-their-own-repository.md) as the nearest survivor) +named `mesh-sdk` "contracts shared across tiers: +types, not behaviour." That line is superseded here, because it draws the boundary in the wrong +place. The boundary that matters is not *types versus behaviour* — it is **how often the thing +changes**. + +The current SDK is the cautionary tale, and its failure is precise. `hal/sdk` holds all the +code, including a per-module API client for every service (`clients/plex.ts`, `clients/gitea.ts`, +…) and a per-module tool implementation for each (`tools/plex.ts`, …). Every module depends on +the SDK, so **every edit to any of that per-module code rebuilds every module** — the cascade. +The SDK is under constant maintenance precisely because it became the place all the volatile +per-module logic accumulated. + +The root cause is worth stating exactly, because the fix follows from it: the pressure was never +to share a client *between* modules. It was to share a client between one module's *own features* +— plex's tools, its health check and its hooks all wanted the same `PlexClient` — and the only +place to share code across a module's features was the global SDK. So **intra-module sharing +leaked out as inter-module coupling.** + +## Decision + +The SDK holds the **stable spine** that modules build against, and earns its place by rarely +changing. The test for membership is change-frequency, not kind. + +### What it holds + +- The **tool-serving harness** — the worker and registration mechanism, and the tool-definition + type. *How* a tool is declared and served is settled; it does not change when an individual + tool does. +- The **messaging and event framework** — the broker client, the event consumer, the envelope. +- The **contracts** — the manifest, declaration, provision and link shapes. +- **Core primitives** — sealing and crypto, semver, the shared resolution helpers. + +These change rarely and deliberately. When one of them does change, a rebuild of everything is +the *correct* outcome, because the contract every module shares has genuinely changed. + +### What it must not hold — the more important half + +- **A module's API client.** A Plex client, a Gitea client, a MinIO client belong in their + module. They change when that service's API or the module's use of it changes, which is often, + and which has nothing to do with any other module. +- **A module's tool implementations.** Same reason, same place: in the module. +- **Anything volatile** — anything that changes when one service's features change. + +The rule, stated so it can be applied without re-deriving it: + +> If editing a thing recompiles unrelated modules **and** it changes often, it does not belong +> in the SDK. + +Both conditions are load-bearing. A rare change that cascades is fine — that is a contract, and +the cascade is correct. A frequent change that stays local is fine — that is a module minding its +own business. Only **frequent *and* cascading** is the disease, and per-module clients and tools +are its carriers. + +### Where per-module shared code lives instead + +Code shared among a module's *own* features lives **in the module**. The default is the plainest +thing that works: an ordinary shared file the features import — `plex/client.ts`, imported by +`plex/tools/`. Within one module, features are files importing sibling files; no package +boundary, no ceremony. + +A **module-local SDK** (a sub-package with its own version) is warranted only for the few modules +whose shared surface is large enough to version on its own. It is the exception, not the shape. + +Either form gives the property the global SDK could not: editing a module's shared code rebuilds +**that module and nothing else**. + +## Consequences + +- The cascade becomes **structurally impossible for module logic**. There is no longer an edge + from one module's internals to another, so the only thing that can rebuild everything is a real + change to a shared contract in the SDK — which is rare, and when it happens, is right. +- The SDK is small and stable **by construction**, not by discipline. Its size is no longer a + thing anyone has to police. +- **Converting a module from the current system is partly a de-coupling, not just a move.** Its + client and its tools are pulled *out* of the shared SDK and *into* the module. A conversion + that copied `clients/plex.ts` into the SDK's replacement would rebuild the exact mistake. +- The host still does not import the SDK. It depends on nothing + ([ADR 0005](0005-the-node-host.md)) and **mirrors** the contracts rather than + importing them, exactly as its apply-shapes table already does deliberately. The SDK is shared + by the tiers that *can* share code; the host is not one of them. + +## References + +- The earlier "repository structure" decision — named the repositories; its `mesh-sdk` + description ("types, not behaviour") is superseded by this record (folded in consolidation; + nearest survivor [ADR 0015](0015-applications-live-in-their-own-repository.md)). +- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the decomposition this serves: code + belongs to the boundary that owns it. +- [ADR 0005](0005-the-node-host.md) — why the host mirrors the contracts instead of + importing the SDK. diff --git a/02-DECISIONS/0040-what-a-module-is.md b/02-DECISIONS/0040-what-a-module-is.md new file mode 100644 index 0000000..c94825d --- /dev/null +++ b/02-DECISIONS/0040-what-a-module-is.md @@ -0,0 +1,101 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-03 +deciders: jochen +reconstructed: false +extends: 0009-modules-and-the-graph.md +--- + +# 40. What a module is + +_Reconciliation note (2026-09-05): supersedes the earlier "grouped by domain" decision, which the consolidation folded into how-we-build.md; no standalone record remains to point at._ + +## Context + +[ADR 0009](0009-modules-and-the-graph.md) settled that everything is a module, but never said what a +module *is* beyond "a directory the mesh processes." That gap let the catalogue's breadth read as a +smell: a module can carry a container, a built image, tools, a provisioner, migrations, health, +config, seat claims, requires and provides — so much that the unit seemed ill-defined. +The earlier "grouped by domain" decision (folded in consolidation; see the reconciliation note +above) tried to organise modules by domain, which is the wrong axis. This record states what a module is, drawn from the cases that +stress-tested it: the shell, i3-vs-sway, umami, and "database." + +## Decision + +**A module is one self-contained piece of software the mesh installs and manages** — everything +needed to make that one thing real and integrable: what runs, the seats it claims, what it provides +to other modules, what it requires from them, and what operates it. + +The **software is the module's identity.** Capabilities, seats and provisioned resources are the +**relationships *between* modules**, not what a module is — and that is what binds a module into one +thing. umami is bound by *being umami*: its container runs umami, its provisioner creates umami sites, +its tools query umami, its `requires` gets umami a database. Every feature serves the one software. + +### The three relationships + +1. **Shared seat** — several modules fulfil a capability and coexist; one may be default. bash, zsh + and fish all join `shell`. +2. **Exclusive seat** — modules contend for a single slot; one holds it. i3 (needs x11) and sway + (needs wayland) contend for `display-session`. +3. **Provide / require** — a provider ships the **provisioner** that creates instances of the + resource it offers and returns sealed credentials; a consumer requires it and the mesh wires the + credential in. Symmetric: umami requires a database *and* provides analytics. + +### Interfaces are mesh-owned; providers adapt to them + +The mesh **defines the interface** for a capability — the provider-neutral contract of what a +consumer receives and how it integrates. Both sides conform: a provider's provisioner **adapts** its +software's real API to the mesh contract; a consumer depends on the **interface**, never on a +provider. Swap one provider for another and the consumer does not change. + +### The naming rule — draw the interface at the consumer's real coupling + +Name a `provides`/`requires` at the **widest boundary across which the consumer genuinely does not +care which implementation serves it**: + +- Where the consumer's coupling is thin — an analytics embed snippet and dashboard, opaque to it — + the mesh defines a neutral interface (`analytics`) and providers (umami, amumi) adapt. Swappable + across vendors. +- Where the consumer **speaks a protocol** — a database's wire protocol and query dialect — the + interface *is* the protocol: `postgres-database`, `mssql-database`, `mongodb-database`. Swappable + only among protocol-compatible implementations, **never across**, because the application cannot + cross it either. "database" is not a capability; the protocol is. +- **Never false genericity.** A name must not promise a swap the contract cannot deliver + ([research 005](../01-RESEARCH/005-domain-grouping/analysis.md)). + +This is [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)'s rule made general — "names the protocol, not +the product; a database names the engine because the app targets it" — with the reason stated: the +contract sits where the coupling is. + +### What is not a module + +- A **library** (built against, never deployed — [ADR 0039](0039-what-the-sdk-holds-and-refuses.md)). +- A **control-plane context** (the mesh itself — [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)). + +A **swappable machine mechanism** (a firewall — ufw, nftables) *is* a module implementing a +capability. The host hardcodes no firewall, supervisor, package manager or runtime; it owns only the +generic apply primitives and platform detection, so it runs where none of those exist — an Android +phone has no ufw, systemd, pacman or Docker. + +## Consequences + +- **Supersedes the earlier "grouped by domain" decision** (folded in consolidation; see the + reconciliation note above). Modules are + organised by their relationships (seats, provisions), not grouped into domain folders. +- **Refines [ADR 0009](0009-modules-and-the-graph.md).** Everything the mesh runs and integrates is + a module — but a module is defined by the *software it delivers*, not by being a bucket of features. +- The target is a **self-fulfilling mesh**: declared wants bound to swappable modules, provisioners + wiring credentials, nothing hardcoded. The control plane's whole job is the binding. +- Converting a module from the old system includes pulling its per-module code out of the shared SDK + ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) and shipping its provisioner as an adapter to a + mesh interface — a de-coupling, not just a move. + +## References + +- [ADR 0009](0009-modules-and-the-graph.md) — everything is a module; this says what one is. +- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — contexts are the mesh, not modules. +- The earlier "grouped by domain" decision — superseded (folded in consolidation; see the note above). +- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — protocol-not-product, generalised here. +- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — per-module code lives in the module. +- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph of these relationships. diff --git a/02-DECISIONS/0041-events-are-a-relationship.md b/02-DECISIONS/0041-events-are-a-relationship.md new file mode 100644 index 0000000..a381307 --- /dev/null +++ b/02-DECISIONS/0041-events-are-a-relationship.md @@ -0,0 +1,85 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-03 +deciders: jochen +reconstructed: false +extends: 0040-what-a-module-is.md +--- + +# 41. Events are a relationship, the lighter sibling of provisioning + +## Context + +[ADR 0040](0040-what-a-module-is.md) names two relationships between modules — seats and +provide/require (provisioning). A third is latent in the mesh and worth making first-class: the +broker every node already runs ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) can carry a +module's activity as **events**, which any other module reacts to. A logger that writes an audit +trail, a module that acts when another module acts, observability — all of it is one mechanism, and +today it is ambient rather than declared. + +## Decision + +**A module emits events and consumes events, and both are declared** — parallel to `provides` / +`requires`, so the mesh knows the event graph the same way it knows the provisioning graph. + +### Events are provisioning's lighter sibling + +| | provisioning | events | +|---|---|---| +| shape | **1:1**, a provider creates a resource *for* one consumer | **1:many**, a module emits, any number listen | +| credential | yes — sealed, per consumer | none — it is broadcast | +| machinery | a provisioner (the reconcile adapter) | nothing but the broker's topic routing | +| declared as | `provides` / `requires` | `emits` / `consumes` | + +Because an event is broadcast and credential-free, there is no provisioner and no per-consumer +setup — only a subscription. That is why it is the *lighter* relationship, and why most +inter-module reaction should be an event, not a provision. + +### An event carries what an audit needs + +Every event carries its **type** (a dotted topic key, so listeners match by prefix), its **source** +module, the **node** it came from, and the **time**. A body follows. The metadata is not optional: +a reaction may only need the body, but an audit trail needs to know who did what, where and when, +and an event that cannot answer that is not auditable. + +### The audit logger is just a consumer of everything + +A logger that records the whole mesh's activity is **not a privileged component** — it is an +ordinary module that consumes `#` (every event) and writes them down. It holds no special access; +it only listens widely. That it falls out of the model with no new machinery is the check that the +model is right. + +### `consumes` is validated like `requires` + +A `consumes` for an event that **nothing** `emits` is a dangling edge, and the mesh refuses it +before deploy — the same rule that catches a `requires` for a resource nothing provides +([research 011](../01-RESEARCH/011-the-module-graph/00-overview.md)). A listener waiting for an +event that can never arrive is a silent failure, and this repository's whole discipline is against +silent failure. + +### One runtime serves all three + +The per-node module runtime that serves a module's tools also wires its `consumes` (subscribe, +dispatch to the handler) and lets its code `emit`. Tools are *invoked* (request/reply), resources +are *provisioned* (1:1, credentialed), events are *emitted and consumed* (1:many, broadcast) — +three relationships, one broker, one runtime, all declared on the manifest. + +## Consequences + +- The mesh gains a declared **event graph** alongside the provisioning graph — visible, validated, + reasoned over. +- **Reaction becomes the default coordination**: a module acts on another's event without either + knowing the other, and without a credentialed link. Coupling drops. +- An **audit trail** is a module, not a platform feature — and can be swapped, extended or run more + than once (a file logger and a queryable one) with no change to anything that emits. +- The runtime must dispatch a module's event handlers as well as its tools; that generalisation is + small (both arrive by importing the module's entrypoint) but it is real work. + +## References + +- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker events ride. +- [ADR 0040](0040-what-a-module-is.md) — the relationships this extends. +- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the + broker binding and the runtime are not. +- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph these edges join. diff --git a/02-DECISIONS/0041-the-host-depends-on-nothing.md b/02-DECISIONS/0041-the-host-depends-on-nothing.md deleted file mode 100644 index ac27923..0000000 --- a/02-DECISIONS/0041-the-host-depends-on-nothing.md +++ /dev/null @@ -1,78 +0,0 @@ ---- -status: accepted -date: 2026-08-26 -deciders: jochen -reconstructed: false -extends: 0037-the-host-applies-it-does-not-decide.md ---- - -# 41. The host depends on nothing that must be installed first - -## Context - -[ADR 0030](0030-the-repository-structure.md) calls tier 0 *"the one binary installed by hand"*, -and [research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) states the property the -whole tier rests on: *"a binary whose whole argument is that it has no dependencies"*. - -Building it forced the question that phrase had been carrying unexamined. Everything else in -the mesh is TypeScript, and `how-we-build` §8 says so. A TypeScript host needs a runtime present -before it can run — so the thing installed by hand becomes **two** things, and the second must -be installed by the means the host exists to replace. - -## Considered options - -1. **TypeScript, with a runtime installed first.** Simplest, and matches every other - repository. Rejected: it breaks the property the tier is built on. A host that cannot run - until something else has been installed by hand is not the bottom of the stack. -2. **TypeScript, bundled as a single executable.** Preserves the language and produces one - file. Rejected on two grounds: it carries roughly ninety megabytes of runtime to preserve a - language choice, and single-executable bundling is a young feature — tier 0 is the worst - place in the system to discover its edges. -3. **A statically linked binary in a language built for it.** Chosen; Go. - -## Decision - -**The host is a single statically linked binary that requires nothing to be present.** Copy it -onto a machine and run it. That is the whole installation. - -**It is written in Go.** The job is system-level — run commands, write files, speak to the -firewall, the overlay, the service manager and the package manager — which is what Go's -ecosystem is for, and it cross-compiles to every architecture the mesh might reach, including -the lighter devices requirement 6 anticipates. - -**The second language costs less here than anywhere else it could appear**, and the reason is -architectural rather than convenient. [ADR 0037](0037-the-host-applies-it-does-not-decide.md) -means the host never queries the mesh database. -[ADR 0039](0039-the-link-is-the-security-boundary.md) means it only ever receives declarations. -So the host shares **no code** with any other tier — not a client, not a schema, not the SDK. -It is joined to the mesh by a message contract and nothing else. - -The language boundary therefore falls exactly on an architectural boundary that already exists. -A second language usually costs duplicated logic; here there is none to duplicate. - -## Consequences - -- **`how-we-build` §8 needs a scope.** It reads *"TypeScript throughout"*, which was true when - everything was a service or a surface. It is now scoped to those, with tier 0 named as the - exception and this record as the reason. That is a constitution change, and the sync it owes - is part of it. -- **Agents must write Go to work on the host.** A real cost, and the one genuine argument - against this. It is bounded by the host being the only thing in tier 0 — nothing else in the - mesh acquires a second language because of this. -- **The dependency-direction lint the design calls for gets easier, not harder.** A Go module - cannot accidentally import a TypeScript control-plane client; the boundary is enforced by - there being no path across it. -- **Cross-compilation replaces per-node builds.** The host is built once per architecture and - copied, rather than built on the machine it runs on — which is what makes *"copy it and run - it"* true rather than nearly true. -- **Two toolchains in the lab.** Scenarios that place a host need a Go build available, and the - lab is TypeScript. The binary is built before the scenario runs, not inside it. -- **This is reversible at a cost that will only grow.** It is being taken at the moment the - first line is written, which is the cheapest point it will ever be taken. - -## References - -- [ADR 0030](0030-the-repository-structure.md) — *the one binary installed by hand*. -- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — why the host shares no code. -- [ADR 0039](0039-the-link-is-the-security-boundary.md) — why it receives declarations only. -- [`05-the-node-host.md`](../03-DESIGN/01-to-be/05-the-node-host.md) — the design this serves. diff --git a/02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md b/02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md new file mode 100644 index 0000000..0479407 --- /dev/null +++ b/02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md @@ -0,0 +1,116 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-03 +deciders: jochen +reconstructed: false +extends: 0041-events-are-a-relationship.md +--- + +# 42. The shape of an event on the wire + +## Context + +[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits`/`consumes`, the +graph, the audit logger. It did not say what an event *is* on the broker: the exchanges, the +routing keys, the headers, the queues and their configuration. That shape is a contract every +emitter and consumer conforms to, exactly as [ADR 0010](0010-delivery.md) +is for declarations — and it was being decided ad-hoc in code. This settles it, so the sdk and the +runtime implement one contract and a module never reinvents it. + +## Decision + +### Two exchanges, kept apart + +- **`mesh.events`** — a durable topic exchange. Every event rides it: module, mesh and node. +- **`mesh.rpc`** — a durable topic exchange. Tool invocations (request/reply) ride it. + +Kept separate because RPC is not an event: a `#` subscription on `mesh.events` is then a complete +audit of what happened, with none of the invocation traffic. + +### The routing key is the event type, namespaced by origin + +Dotted and hierarchical — `..` — with three reserved origins: + +- `module..` — `module.umami.site.created` +- `mesh..` — `mesh.delivery.deployed`, `mesh.provisioning.granted` +- `node..` — `node.anchor.joined`, `node.anchor.unreachable` + +Topic matching gives a consumer `node.*.joined`, `module.umami.#`, or `#`. The origin roots are +reserved; everything after is the emitter's own namespace. + +### Metadata in headers, payload in the body + +An event's identity and provenance are AMQP **headers**, so a consumer — or the broker, or an +audit tool — reads who/when/what without parsing the body, and the body is only the domain payload. + +**Required headers** + +| header | meaning | +|---|---| +| `x-event-id` | a unique id — for dedup and audit (delivery is at-least-once, below) | +| `x-source` | the emitter: the module, context or node name | +| `x-node` | the node it was emitted from | +| `x-time` | emit time, RFC-3339 | +| `content-type` | `application/json` | + +**Optional headers** + +| header | meaning | +|---|---| +| `x-causation-id` | the event or command that caused this one — tracing | +| `x-schema` | a version of the body's shape, so a body evolves without silent misreads | + +The routing key already carries the type; it is not duplicated as a header. An **unknown `x-` +header is ignored, not refused** — unlike a declaration, an event is observed by parties that need +not all understand every header, and refusing would couple every consumer to every emitter's +additions. + +### Messages are persistent + +Events are published persistent (delivery-mode 2). An audit trail that loses events on a broker +restart is not one, and the cost is disk the broker already spends on everything durable. + +### Queues: one per consumer, durable, dead-lettered + +- **A consumer's queue** is `..events`, durable, bound to that module's consumed + patterns. Durable so a restart does not drop what arrived while it was down. **Manual ack** after + the handler succeeds — at-least-once. +- **Prefetch** bounds in-flight work (default 32) so one slow consumer does not pull the whole + backlog into memory. +- **A dead-letter exchange** `mesh.events.dead` receives a message rejected past a redelivery limit, + so a poison event is set aside for inspection rather than looping forever or vanishing silently. +- **The audit logger's queue** `.audit-logger.events`, bound to `#`, is the same shape — + durable, persistent, dead-lettered — because completeness is its whole job. +- **RPC reply queues** are exclusive, auto-delete and server-named; **RPC serve queues** + `serve.` are durable and shared, so several runtimes serving one tool key compete rather than + each answer. + +### At-least-once, and consumers are idempotent + +A handler may see an event twice — a redelivery after a crash between doing the work and acking. +Consumers must be idempotent, and `x-event-id` is what makes dedup possible. **Exactly-once is not +offered**: it is a promise no broker keeps honestly, and saying so is better than pretending. + +## Consequences + +- The event shape is a versioned, enforced contract, not conventions each module reinvents. The + sdk's `emit`/`on` and the runtime's AMQP binding implement it; a module never sees an exchange or + queue name. +- Metadata-in-headers means the body is exactly the domain payload, and a consumer that only wants + provenance never parses it. +- Adding a header or an origin root widens the contract and is reviewed as one — the discipline + [ADR 0010](0010-delivery.md) applies to the + declaration vocabulary. +- The sdk's first cut carried source/node/time in the *body*; this supersedes that — they move to + headers. That is code to align, in `mesh-sdk` (`emit`/`on`) and `mesh-tools` (the binding, queue + config, dead-letter). + +## References + +- [ADR 0041](0041-events-are-a-relationship.md) — events as a relationship; this is their wire shape. +- [ADR 0010](0010-delivery.md) — the precedent: a wire + contract, versioned, additions reviewed as security. +- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker. +- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the + binding, queue config and dead-letter are the runtime's, not the sdk's. diff --git a/02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md b/02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md deleted file mode 100644 index d11120d..0000000 --- a/02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md +++ /dev/null @@ -1,140 +0,0 @@ ---- -status: accepted -date: 2026-08-26 -deciders: jochen -reconstructed: false -extends: 0037-the-host-applies-it-does-not-decide.md ---- - -# 43. A declaration is an ordered list of resources the host owns - -## Context - -[`05-the-node-host.md`](../03-DESIGN/01-to-be/05-the-node-host.md) leaves *what a declaration -is* open and calls it the first thing to settle in build. Stage 2 — applying with no mesh -present — cannot start without it. - -Three constraints already bind it, and between them they decide most of the shape: - -- **Data, not instructions**, with a finite, versioned vocabulary and anything outside it - refused rather than interpreted ([ADR 0039](0039-the-link-is-the-security-boundary.md)). -- **The host applies; it does not decide** ([ADR 0037](0037-the-host-applies-it-does-not-decide.md)). -- **The host depends on nothing** ([ADR 0041](0041-the-host-depends-on-nothing.md)). - -## Decision - -### JSON, because the host has no dependencies to spend - -Go's standard library carries `encoding/json` and no YAML. A YAML declaration would put a -third-party parser inside the one binary whose entire argument is that it needs nothing — to -gain authoring comfort in a document that is, in the ordinary case, generated by a machine and -read by a machine. - -The mesh's *authoring* formats stay YAML. What crosses the link is JSON. - -### An ordered list, because ordering is a decision - -A declaration states the order its resources are applied in. The host does not sort, does not -resolve dependencies, and does not decide what must come before what. - -This follows from [ADR 0037](0037-the-host-applies-it-does-not-decide.md) more strictly than it -first appears. A host that derived ordering from declared dependencies would be **deciding**, -and it would be deciding the thing most likely to differ between what the control plane -intended and what the machine does. The control plane knows what depends on what; it says so by -saying when. - -Consequence accepted: the control plane must order correctly, and a mis-ordered declaration -fails at the step that needed something not yet there — which is at least the *right* failure, -naming the resource rather than a mystery. - -### Every resource has a stable identity - -Not a position, not a hash of its content: a name the control plane keeps stable across -declarations. It is what lets the store say *this is the same resource I applied last time*, -which is what makes convergence possible at all. - -### Unknown is refused, never skipped - -An unknown declaration version, an unknown resource type, or an unknown field is a **refusal of -the whole declaration**. Not a warning, not a skip, not best-effort. - -A host that skipped what it did not understand would apply most of a declaration and report -success — a node that looks configured and is not, which is -[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) with the -declaration on the other side of the wire. Refusing whole also means an older host cannot be -handed a newer vocabulary and quietly do half of it. - -### Complete for what the host owns, and only that - -*Desired state* invites the question of removal, and the honest answer needs a boundary. - -**The host removes what it previously applied and is no longer declared.** It knows what it -applied because it recorded it (`store`), so this is a fact it holds rather than an inference. - -**The host never removes anything it did not create.** A machine has things on it that the mesh -did not put there, and a converger that treats *not declared* as *must not exist* deletes them. -The rule that prevents production data loss elsewhere in this repository is the same one: -[ADR 0018](0018-the-mesh-creates-no-symlinks.md) exists because a tool did something to a path -it did not own. - -So: authoritative over its own footprint, inert everywhere else. - -### Addressed, and checked when it can be - -A declaration names who it is for. A host that has an identity refuses one addressed elsewhere. -A host that has no identity yet — the first node, applying the bundle it carries — has nothing -to check against and applies it. - -## Where the list comes from - -This record specifies what the host **accepts**. What produces a declaration is deliberately -not settled here, and the reason is worth stating rather than leaving as an omission. - -**Today, and at stage 2: by hand.** `substrate.lock` is authored and pinned — a person writes -the resources and writes the order. That is the first node's path, where there is no control -plane to derive anything from. - -**Afterwards: the control plane derives it**, from three things it already holds — which -modules are assigned to this node, what those modules' configuration resolves to, and what each -module declares it needs. - -**And the order comes from the graph.** Each module expands to resources; the modules are -ordered by their declared dependencies on one another. That is -[research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — `requires`, `provides`, -`excludes` — and a declaration is the graph's output, flattened for one node. - -So this record is complete on the consumer side and silent on the producer side, because the -producer does not exist and its shape is what 011 is investigating. The consumer can be settled -first because the host must refuse what it does not understand whoever wrote it. - -**What this means for ordering.** [ADR 0037](0037-the-host-applies-it-does-not-decide.md) puts -the ordering decision in the control plane; 011 decides how the control plane makes it. If the -graph turns out not to determine a total order, that is 011's problem to solve and not the -host's — the host will still be handed a list, and will still apply it as given. - -## Consequences - -- **Ordering is now a control-plane responsibility**, and getting it wrong is a class of bug - that will appear. It is the correct place for it: the alternative puts a dependency solver in - tier 0 and a decision in the wrong tier. -- **The vocabulary is a security artefact.** Every type added widens what a compromised control - plane can express, so additions are reviewed as such rather than as features. -- **Removal is bounded but not free.** A resource dropped from a declaration is deleted on the - next apply, so removing a line is an act with an effect — which is the point, and is worth - saying out loud because it does not look like one. -- **The store becomes load-bearing at stage 2**, earlier than the build order suggests. Nothing - can be removed without knowing what was applied, so the record of applied resources arrives - with the first apply rather than with the link. -- **A closed address space bounds what the first types can be.** A scenario has no route to a - package repository, so a declaration whose resources must be fetched cannot be applied in the - lab at all. The first vocabulary is therefore what needs no network — files, directories, - service state — and packages and containers wait on *where `place:` gets its artifacts from*, - which is open in - [`02-scenario-declaration.md`](../03-DESIGN/01-to-be/02-scenario-declaration.md). - -## References - -- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — why ordering is not the host's. -- [ADR 0039](0039-the-link-is-the-security-boundary.md) — bounded by form. -- [ADR 0041](0041-the-host-depends-on-nothing.md) — why JSON. -- [ADR 0008](0008-a-failed-step-fails-the-job.md) — why a refusal is whole. diff --git a/02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md b/02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md new file mode 100644 index 0000000..cefd111 --- /dev/null +++ b/02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md @@ -0,0 +1,101 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-04 +deciders: jochen +reconstructed: false +extends: 0041-events-are-a-relationship.md +--- + +# 43. A module's broker account is scoped by what it emits and consumes + +## Context + +[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits` and `consumes` +on the manifest. [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) gave them a wire shape — the +`mesh.events` exchange, the durable per-consumer queue, the reserved routing-key origins. Neither +said how a module *reaches* the broker: what account it holds, and what that account is allowed to +do. + +As the code stands, there is no answer. The mesh can provision a **node** account (at enrolment) +and a **builder** account (scoped to the build queue), and it can *deliver* any module a sealed +own-secret at a declared path — but it has no way to provision a broker **account** for a general +module. A module that declares `own-secrets: {broker: …}` and nothing more receives thirty-two +random bytes, not a credential. So on the broker, `emits` and `consumes` are enforced by nothing: a +running module could bind any queue, consume any pattern, and publish under any origin, and the +manifest that says otherwise would be describing a boundary no code draws — the exact shape of fault +[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) records, a scope +declared in manifests and read by nothing. + +This settles it, so a module's place on the bus is a thing the broker enforces rather than a thing +the manifest merely claims. + +## Decision + +### A module gets a broker account when it is assigned, and its permissions are the manifest + +When the mesh assigns a module to a node it provisions a broker account for that module on that node, +sealed to the node ([ADR 0004](0004-a-node-and-how-it-joins.md)) and delivered as the +module's `own-secrets` broker — `amqps://` with the mesh's fingerprint, the shape +[ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) already carries. The account's permissions are +derived from the manifest, and are exactly these: + +- **What it consumes.** Read on `mesh.events`, and configure-and-read on its own queue + `..events` bound to the patterns in `consumes`. It cannot bind or read another + module's queue. A module that consumes nothing gets no read on the events exchange at all. +- **What it emits.** Write to `mesh.events`, restricted to routing keys under its own origin, + `module..*`. It cannot publish as another module, and cannot publish under the reserved + `mesh.*` or `node.*` origins — those belong to the mesh and the host (ADR 0042). A module that + emits nothing gets no write. +- **Nothing else.** The events account reaches `mesh.events` and that module's own queue, and no + more. Tool serving and calling over `mesh.rpc` is a separate grant on the same principle — a + module serves the tool keys it declares and calls the ones it is bound to — and is scoped the same + way rather than folded in here. + +### Consuming everything is a privilege, granted deliberately + +`consumes: ["#"]` — the audit logger — is read across the whole bus: every module's events, the +mesh's, every node's. That is not a pattern like any other; it is the power to see everything, and +the account is where it becomes visible. The grant that lets one module read the entire bus is one +the mesh issues on purpose and can be audited — the answer to *who can read everything* is a row, not +a guess — rather than a breadth any manifest acquires by typing a single character. A `#` consume is +a reviewed grant, not a default one. + +### The account is how the declaration is enforced + +Because the account can do only what `emits` and `consumes` name, the broker itself refuses a module +that tries to consume a queue it did not declare or emit under an origin it does not own. That is what +makes an event relationship a rule and not a comment — the discipline that a stated rule says how it +is checked. A manifest that over-declares grants more than the module uses, which is visible and +reviewable; one that under-declares makes the module fail closed at the broker, which is the safe +direction to be wrong in. + +## Consequences + +- The control plane gains a **generic module broker-account**, derived from the manifest. The + builder stops being a special case: its access to the build queue becomes an ordinary expression of + what it consumes and serves, not a bespoke account method. One rule, and the builder is an instance + of it. +- The runtime reads its credential from a file (the broker own-secret), `amqps://` verified against + the mesh's fingerprint. The `guest` account is for raising the substrate, never for a module — a + module documented as holding its own credential and handed the broker's administrative one is worse + than one with no credential story at all. +- `emits` and `consumes` stop being advisory. They are the module's authority on the bus, so the + manifest is now a security boundary and is reviewed as one, the discipline + [ADR 0010](0010-delivery.md) applies to the declaration + vocabulary. +- *Who can read the whole bus* becomes an answerable question, because `#` is a grant and not an + accident. + +## References + +- [ADR 0041](0041-events-are-a-relationship.md) — events are a relationship; this scopes the account + by that relationship. +- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the wire this account secures: the queue, + the origins, the `amqps` credential shape. +- [ADR 0004](0004-a-node-and-how-it-joins.md) — the link is the security boundary; a + module's account is sealed to its node the same way a node's is. +- [ADR 0010](0010-delivery.md) — a declaration is owned + and its additions reviewed; a module's broker permissions are that discipline applied to the bus. +- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a scope declared + in manifests and enforced by no code: the fault this decision closes for events. diff --git a/02-DECISIONS/0044-a-public-name-is-provisioned-like-any-capability.md b/02-DECISIONS/0044-a-public-name-is-provisioned-like-any-capability.md new file mode 100644 index 0000000..d86e0a5 --- /dev/null +++ b/02-DECISIONS/0044-a-public-name-is-provisioned-like-any-capability.md @@ -0,0 +1,94 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-04 +deciders: jochen +reconstructed: false +extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md +--- + +# 44. A public name is provisioned, not registered by hand + +## Context + +The mesh names and resolves its own machines internally: the overlay generates +`..` wildcards, dnsmasq answers them (`wildcard-resolution`), and the mesh +issues a certificate for each internal name. A service reachable at a *public* domain — +`plex.example.com`, not `plex.anchor.internal` — needs three things that machinery does not give it: + +- a **public DNS record** at a registrar or DNS provider, so the name resolves on the internet; +- a **publicly-trusted certificate** for it, because the mesh's own authority is trusted by nobody + outside the mesh; +- and routing from that name to the module — which the reverse proxy already does: a module + `requires` the `route` capability and the proxy provides it, routing by the host it was asked for. + +The routing exists. The public DNS record does not: the mesh has no way to make a name resolve on +the public internet, so today that is a step someone does by hand at a DNS provider, outside the +mesh, remembered nowhere. A public name is therefore the one part of reaching a service that the +declaration graph cannot grant or withdraw — which means it is created once and outlives whatever it +was for, the shape of drift this project exists to remove. + +## Decision + +### A public name is a capability, requested like any other + +A module reachable at a public host declares `requires: ["public-dns"]` and contributes the hostname +it wants — beside `requires: ["route"]`, which exposes it through the proxy. The name is then +provisioned on declaration ([ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)): created +when the module is assigned, removed when it is withdrawn, reconciled like every provision. + +### The interface is neutral; the providers are the registrars + +`public-dns` is drawn at the consumer's coupling: the consumer wants *a public name that resolves to +me*, and does not care whether Cloudflare, Route 53 or a registrar's own API puts the record there. +So the interface is neutral and the providers are provider-scoped — `cloudflare-dns`, +`route53-dns`, `porkbun-dns` — each implementing the one `public-dns` contract, the same way a +neutral database coupling is answered by `postgres-database` and `mssql-database`. A module names +`public-dns`; it never names a registrar. + +### The record points at the mesh's public ingress, not at the node + +What the name resolves to is the address the reverse proxy answers on, not the consuming machine's. +A public service is reachable only *through* the proxy — the proxy holds the `route` grant and routes +by host to the module — so the public name must resolve to the proxy. `public-dns` and `route` are +the two halves of one public exposure: the name, and what the name reaches. + +### The record is a fact, not a secret + +A DNS record is public by definition, so the grant returns the fully-qualified name and its TTL and +nothing sealed. The only secret is the provider's own API credential, which is the provider module's +own-secret and never leaves it — the module that wanted the name never sees it. + +### Events + +The provider emits `module..record.created` and `module..record.removed` +([ADR 0041](0041-events-are-a-relationship.md)), so *which names the mesh publishes, and where* is a +question answered from the event trail and the grants, not from a folder of records edited at a +provider. + +### The public certificate is the proxy's, and is named here only to pair it + +A public name without a publicly-trusted certificate is reachable and not trusted — the same pairing +the internal name and the mesh-issued certificate already have. Obtaining that certificate (ACME +against the now-resolving public name) is the reverse proxy's to do, and its mechanism is its own +decision; it is named here so the pairing is not forgotten, not resolved here. + +## Consequences + +- A public name is created and torn down with the module, so it cannot outlive it, and the mesh can + say which public names it publishes without anyone reading a registrar's dashboard. +- Adding a registrar is adding a provider that answers `public-dns`; the modules that want names do + not change. +- Public exposure of a service is a trio of separate, declared, enforced relationships: the firewall + opens the proxy's public port ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)), + `route` routes the host to the module, and `public-dns` makes the host resolve. + +## References + +- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — a capability is provisioned on + declaration; a public name is one. +- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall, the other + half of the reachability question this was asked with. +- [ADR 0041](0041-events-are-a-relationship.md) — the provider's record events. +- [ADR 0040](0040-what-a-module-is.md) — a provider and its interface; the neutral-interface, + scoped-provider naming this follows. diff --git a/02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md b/02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md new file mode 100644 index 0000000..a5c61ac --- /dev/null +++ b/02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md @@ -0,0 +1,93 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-04 +deciders: jochen +reconstructed: false +extends: 0005-the-node-host.md +--- + +# 45. A machine's firewall is the sum of what its modules listen on + +## Context + +The reverse proxy is a *provider*: a module `requires` the `route` capability and a running proxy +provides it, routing traffic by name and reaching back to the consumer. A fair question follows — +is the firewall the same shape? Should a module *register* a port with a firewall provider the way +it requests a route? + +It should not, and the difference is the point. A reverse proxy is a service another component +performs; a firewall is a property of the machine — a packet filter the host applies to itself. +Modelling it as a provider would invent a credential and a reach-back for something that has neither. + +And the mesh already has the registration: a module declares `listens: [{ port, from }]` — the port +it accepts connections on, and from where. That *is* how a service says it wants a port open. What is +missing is not a model but enforcement. [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) +records that a `scope:` key five manifests carry is read by no code: a manifest can appear to +restrict a port and restrict nothing — the exact fault +[how-we-build.md](../00-META/how-we-build.md) names, *an unenforced rule is indistinguishable from a +wrong one*, made worse because the declaration reads as a restriction. + +## Decision + +### The firewall is derived and host-applied, not a provider + +A machine's firewall is the sum of what the modules assigned to it declare they listen on, computed +by the host and applied as one of its owned resources ([ADR 0005](0005-the-node-host.md): +the host applies, it does not decide; [ADR 0010](0010-delivery.md): +the declaration is owned resources). It is not a capability, not a per-consumer grant — opening a +port is a declarative fact about a machine, so it is computed and applied, not requested and +credentialed. + +### `from` is the whole of public-versus-internal + +The distinction the question is really about lives in `from`: + +- `listens: [{ port: 5432, from: mesh }]` — open to the private overlay only. +- `listens: [{ port: 443, from: anywhere }]` — open to the public internet. + +A module registers a port on the firewall by listening on it and saying from where. There is no +separate firewall capability, because the firewall is not a thing that reaches back or holds a +secret; it is the machine's own filter over the ports its modules named. + +### The host enforces it both ways, and unknown keys are refused + +A port a module listens on is opened to exactly the scope it named; a port nothing declares is +closed. And a key the firewall does not read — the `scope:` of issue 003 — is refused at the +manifest, not accepted and ignored, so a declaration that reads as a restriction is one. This is the +discipline [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) applied to the +broker account, applied here to the packet filter: the declaration is the enforcement, or it is a +comment. + +### A public service is exposed through the proxy, not by opening its own port + +Reaching the public internet is normally not `from: anywhere` on the service's own port. The service +listens `from: mesh` — only the proxy reaches it — and `requires: route`, so the sole machine with a +public opening is the one running the reverse proxy, and the service is exposed by name through it. +`from: anywhere` is the deliberate direct-exposure case, for a service that is its own front door. + +## Consequences + +- Issue 003 is closed: the firewall is computed from `listens` and enforced, so a declared scope is + real and an undeclared port is shut. Rejecting unknown manifest keys is the general fix, of which + the `scope:` key was one instance. +- The firewall and the reverse proxy stop being confused for one model: the firewall is the machine's + filter (host-derived from `listens.from`); `route` is a name-router (a provider); the public DNS + name is a third thing ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)). A + public service uses all three. +- The modelling question is answered: a module registers a port by declaring `listens`, and reaches + the public internet by name through `route` + `public-dns` — never by the firewall being a + provider. + +## References + +- [ADR 0005](0005-the-node-host.md) — the host applies; the firewall is one of + the things it applies. +- [ADR 0010](0010-delivery.md) — the firewall is a derived + owned resource, not a grant. +- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the same discipline: + a declaration is enforced, or it is a comment. +- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the public name, the other + half of the reachability question this was asked with. +- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the unenforced + `scope:` this closes. diff --git a/02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md b/02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md new file mode 100644 index 0000000..402c5ba --- /dev/null +++ b/02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md @@ -0,0 +1,88 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-04 +deciders: jochen +reconstructed: false +extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md +--- + +# 46. A module's configuration is its assignment's, not its manifest's + +## Context + +A module is assigned to a node — `assign `, always to a machine; there is no +assignment to the mesh. "Mesh" is a *scope*, not a place: a `provides` or a `claim` scoped `mesh` +reaches the whole mesh, but the module still runs on a node. So the two kinds of thing a module can +carry are the manifest (what the module *is*) and, separately, what it should do *here* — which +differs by deployment and by node. + +The mesh already has the second: **settings**. `settings set [--node ]` — with a node +it is that machine's, without it the whole mesh's — layered over what the module declares and applied +at resolution, changeable without editing the module and without a rebuild. That is the surface a +meshboard would edit. + +But settings today reach only a module's **config-file content** (a mergeable file the module owns). +Configuration that is not a file has been landing in the manifest instead, statically — a registrar's +zone and domain, the address public names point at, and, most sharply, `listens.from`. That last one +is the tell: whether a port is open to the private overlay or to the public internet is a +*per-node deployment choice* — the same database internal on one machine and public on another — and +a value fixed in the manifest is one value for every machine, so it cannot be. Static configuration in +the manifest is configuration in the wrong place: it cannot vary per node, and it cannot change +without a new module version. + +## Decision + +### The manifest is identity and defaults; the assignment's settings are the configuration + +A module's manifest declares what it is — what it provides, requires and claims, the shape of its +resources — and, for anything configurable, a **default**. The values that make a running instance +*this* instance are settings, carried by the assignment: per-node, or mesh-wide when no node is named, +applied over the defaults at resolution. Change one and the next reconcile carries it; nothing is +edited on a machine and nothing is rebuilt. + +### Settings drive the configurable fields the manifest marks, not only file content + +Settings extend beyond a config file's content to the manifest fields a module declares settable — +foremost: + +- **`listens.from`**: a module declares its safe default (`from: mesh`), and a per-node setting + raises or lowers it. postgres declares `listens: [{ port: 5432, from: mesh }]`; on the machine that + should expose it, a setting makes that port `from: anywhere`. Same module, different exposure, and + the firewall ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)) is computed + from the effective value, so the packet filter follows the setting. +- **A provider's own configuration**: a registrar's zone, domain and the ingress its names point at + ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)) are mesh-wide settings, not + manifest constants — one mesh's Cloudflare zone is not another's, and the module description is the + same for both. + +### Unset is the default, and an unknown setting is refused + +A field with no setting keeps the manifest's default, so a module runs correctly configured by nobody. +A setting that matches no settable field — like a config value that reaches no file today — is named, +not silently dropped, so a misspelled setting is found rather than believed (the discipline of +`UnusedSettings`, and of [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md): +a declaration is enforced or it is a comment). + +## Consequences + +- The postgres case works: one module, `from: mesh` by default, `from: anywhere` where a setting says + so — internal on ace, public on novox, changeable live. +- Provider modules stop carrying a mesh's specifics: `cloudflare-dns` describes *a Cloudflare + registrar*, and *which* zone and ingress is a setting, so the same module serves every mesh. +- Configuration becomes a thing a meshboard manages — set per node or mesh-wide, applied on the next + reconcile — rather than a manifest edit and a rebuild ([ADR 0011](0011-managed-files-are-generated-never-edited.md): + the way you change a managed thing is not by editing it). +- What a manifest may not do is grow a value that differs per machine; if it differs per machine it is + a setting, and the manifest holds only the default. + +## References + +- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — what is provisioned on + declaration; its per-instance values are the assignment's. +- [ADR 0011](0011-managed-files-are-generated-never-edited.md) — a managed thing is changed through + the mesh, not by editing it; settings are that, for configuration. +- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall follows the + effective `listens.from`, so making `from` a setting makes exposure a setting. +- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the provider whose zone and + ingress are settings, not manifest constants. diff --git a/02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md b/02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md new file mode 100644 index 0000000..1ed3edd --- /dev/null +++ b/02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md @@ -0,0 +1,87 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-04 +deciders: jochen +reconstructed: false +extends: 0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md +--- + +# 47. A module runs its code as its own process, with its own account + +## Context + +A module is one self-contained thing ([ADR 0040](0040-what-a-module-is.md)), and it gets a broker +account scoped to what it emits and consumes ([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)). +The catalogue now gives modules **tools** and **events** — real code, in the module ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) — +but nothing has said what *runs* that code. The audit-logger showed one shape and was treated as an +exception: a container running the tool runtime carrying the module's compiled code, holding the +module's own scoped account. Every module with tools or events needs the same, and the tempting +alternative does not work. + +**A node-wide runtime that loaded every assigned module's code cannot hold a per-module account.** It +would run under one account with the union of every module's permissions — able to emit as any of +them and read any of their queues — which is exactly the isolation ADR 0043 exists to draw. So the +runtime is per-module, not per-node, and treating the audit-logger as special left the other +modules' code with nothing to run it: the conversion produced tools and events that, as it stands, +never execute. + +## Decision + +### A module with tools or events runs a process of its own + +A module that has tools or events runs a **runtime process** — a container, the tool runtime carrying +that module's compiled code — assigned and started like the module it is, holding the single broker +account the mesh scoped to it (ADR 0043). One module, one process, one account. + +### It serves its tools, each on its own key + +A tool is served on its own key (`serve.`), and a caller invokes a named tool. Only the module +that serves it answers, and the module's account is scoped to exactly its tool keys — so one module +cannot answer another's calls, the isolation ADR 0043 gives events extended to tools. This supersedes +a single `tools.invoke` endpoint that dispatched by name: that shape assumed one runtime for the +whole node, and per-module runtimes competing on one key would each be handed calls for tools they do +not have. + +### It runs its events in the same process, under the same account + +Emitting under the module's own origin and consuming its own queue ([ADR 0042](0042-the-shape-of-an-event-on-the-wire.md)) +happen in that same process, with that same account — not a second one to scope and seal. A module's +tool code, its event code and, for a provider, its provisioner are the one module's code and run as +the one module's process. + +### The runtime image is the tool runtime plus the module's code + +Built from the module's source like any module image — the audit-logger's shape, made the rule, not +the exception. The module declares a `container` for it carrying `MESH_BROKER_FILE` (its sealed +credential, ADR 0043) and its compiled code. A module with **neither** tools nor events runs no such +process: a plain service module — the plex *server*, dnsmasq the resolver — is its service and files +and nothing more. A module that is both a service and code declares both containers: the service, and +the runtime beside it. + +## Consequences + +- The catalogue's tools and events become runnable: each tools-or-events module gains a runtime + container with its scoped credential, and the audit-logger stops being special. Until this, the + converted modules held code with nothing to execute it. +- A process, and a small image, per tools-or-events module. That is the cost of ADR 0043's isolation: + one account per module means one process per module. It is paid deliberately — a shared runtime is + cheaper and cannot be scoped, and a mesh where any module can emit as any other is not one worth the + saving. +- `serve.` per key replaces the single `tools.invoke` dispatch. The sdk's serving and a module's + account scope both come to name tools individually. +- **A provider's provisioner is a runtime process too.** It already runs as its own container; its + events (`bucket.created`, `database.provisioned`) belong to *that* process and need the same + credential. So a provisioner that emits carries `MESH_BROKER_FILE` and its scoped account like any + runtime — or it does not emit. (This is the fix for provisioners that emit today with no broker + bound: the emit is a runtime's, and the provisioner is a runtime.) + +## References + +- [ADR 0040](0040-what-a-module-is.md) — a module is one self-contained thing; its code runs as one + process. +- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the scoped account + this process holds, and the isolation that makes it per-module. +- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the events this process runs, and the + `serve.` queue tools now use. +- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — the code lives in the module; this runs it. diff --git a/02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md b/02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md new file mode 100644 index 0000000..e047950 --- /dev/null +++ b/02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md @@ -0,0 +1,131 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-05 +deciders: jochen +reconstructed: false +--- + +# 48. A provider creates the credential the mesh minted, and seals nothing + +## Context + +A provider module stands up a per-consumer resource — a database, a cache bucket, an object +store user — and the consumer must end up holding a credential that authenticates against it. +Building the module runtime (ADR 0047: a module runs its own code as its own process under its +own account), the provider's provisioner was run for the first time as a delivered thing, and +it did not work. It reads a seal key from the environment that nothing sets, and it seals every +credential it produces to that key with a symmetric passphrase. + +Tracing the credential's path turned up something larger than a missing key. **The provisioner +harness the whole catalogue is built on describes a credential flow the mesh does not have, and +duplicates — incorrectly — one it does.** + +What the sdk's `runProvisioner` does today: + +- reads request files named `*.grant.json` — which nothing in the mesh writes; +- calls an adapter whose `create` **generates its own password** and returns it; +- seals that password with a symmetric key (`$MESH_SEAL_KEY`) and writes a `*.credential` + file — which nothing in the mesh reads, and no consumer ever unseals. + +What the mesh already does, and has wired end to end: + +- The control plane mints one password per (consumer, provider) pair (`Inventory.SecretFor` → + `secrets.Make`) and seals it to **both** node keys asymmetrically — a copy the consumer's + host can open and a copy the provider's host can open. No shared symmetric key exists + anywhere, on purpose: a key both ends hold is a key the mesh would have to distribute, which + is the same problem one level down, and the control plane's own code refuses it. +- The provider is handed, at the path its `receives` names, one contribution per consumer: + the **login to create** (`As`, derived by the mesh so the two ends agree by construction), + the consumer's address and requested values, and a **`Secret` file** holding that consumer's + password sealed to the provider and unsealed onto the machine by its host. +- The consumer is handed the *same* password, as plaintext its own host wrote by unsealing its + copy and substituting it into a config file. The consumer never unseals anything itself and + holds no key. + +So the password a provider's provisioner invents is not even the password the consumer was +given: a consumer authenticating with the mesh's password against a resource the provisioner +created with its own would simply fail. The symmetric seal is not an incomplete feature to +finish delivering a key for. It is a second, contradictory credential model bolted beside the +real one, and it cannot be made to work without building the very thing the mesh was designed +not to have. + +This is a decision and not a patch because the harness is the **provider contract**. Every +provider — the four that exist and the many a real mesh grows — is built on +`runProvisioner(resource, adapter)`. Whatever it says a provider is, they all inherit; and +changing it later is one migration per provider. It is cheaper and more honest to settle what +a provider is now. + +## Decision + +**A provider is handed the credential; it does not make one, does not seal one, and does not +hand one back.** The provisioner's only job is to make the mesh's grants true in its own +software. + +Concretely, for the sdk harness and the adapter contract: + +- The harness reconciles the **contributions the mesh delivers** to the provider's `receives` + path — the list of consumers, each with its login name (`As`), address, requested values, + and the path to its unsealed password (`Secret`). It does not read `*.grant.json` and it + does not write `*.credential`. +- For each consumer present, the harness reads the password from that consumer's `Secret` file + and calls the adapter to bring the resource into being under the given login. For each + consumer no longer present — the mesh drops it from the contributions file when its consumer + goes away — the harness calls the adapter to withdraw it. +- The adapter shrinks to the per-software half and nothing else. It is given the login, the + password, and the values, and it makes the resource exist or removes it. It generates no + password, derives no name, seals nothing, and returns no credential: + roughly `create({ as, password, values })` and `remove({ as })`, both returning nothing. +- `$MESH_SEAL_KEY`, the symmetric `seal()`/`writeSealedCredential` path, and the `*.grant.json` + / `*.credential` files are removed from the provisioning path entirely. The credential + reaches the consumer through the mesh's own asymmetric channel, which already crosses node + boundaries and holds no shared secret. + +Identity stays the mesh's to say. The login the provider creates is the name the mesh derived +and gave the consumer to present; the provider never invents a name, because a name the +consumer cannot learn is a name it cannot authenticate with. + +## Consequences + +- A provider module becomes smaller and unable to be wrong in this way: with no password to + generate and no key to seal to, the class of bug where the two ends hold different secrets + cannot be written. A provider added after this inherits the corrected contract and has no + seal to reintroduce. +- The four current providers (redis, postgres, minio, umami) each lose their `generatePassword` + + seal code and gain a `create` that takes the password it is given. Their teardown becomes + "withdraw the login named `As`". +- The symmetric `seal()`/`unseal()` primitive loses its only caller and leaves — checked, not + assumed: nothing else in the sdk or the catalogue called it, so it is removed with the + provisioner it belonged to. +- **How this is verified:** redis is assigned as a provider in the lab, the contributions and + the unsealed password the mesh would deliver are put in its `receives` path, and a client + authenticates as that consumer with the mesh's password and gets PONG — where a provider that + invented its own password answers WRONGPASS — with `$MESH_SEAL_KEY` set nowhere and no + `.credential` file written. Proven: `provider-uses-mesh-credential` is green. + +**What this does not cover — credential provisions, not data provisions.** This decision is about a +provision whose credential is a *secret the mesh mints* — a login and password (redis, postgres, +minio). A provider that instead *generates* the thing the consumer needs, and that thing is not a +secret — umami's `analytics`, where the consumer wants back a `siteId` umami assigned — does not fit, +because a contract that returns nothing has no way to hand that data back. The seal-key fault was +never umami's (it sealed no password; it returned a public id), so removing the seal does not break +it further, and it still reconciles its sites off the mesh's contributions. But delivering +provider-generated data back to a consumer is a *return path* the mesh does not have and this +decision does not build — a separate shape, left to a separate decision. +- Teardown beyond "remove the login" — data an object store leaves behind when a consumer + leaves — is named by each provider's adapter, not by the harness, and is out of scope here + except to say the contract must leave room for it. + +## References + +- [04-ISSUES/032](../04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md) — the + observation and the cross-repo trace this decision rests on. +- ADR 0047 (the module runtime) — what first ran a provider's provisioner as a delivered + process and exposed this; link to be filled when 0047 lands on the trunk. +- Control-plane mechanisms this relies on already existing: `mesh-control` — + `internal/inventory/secrets.go` (`SecretFor`, `SecretsFrom`), `internal/secrets/seal.go` + (`Make`, the two-blob asymmetric sealing), `cmd/mesh-control/plan.go` (`grantsFor`, the + `Grant.Sealed = ForProvider` delivery), `internal/catalogue/declaration.go` (the `receives` + contribution: `As`, `At`, `Values`, `Secret`). +- The path being removed: `mesh-sdk` — `src/provisioner/index.ts` (`runProvisioner`, `sealKey`, + `writeSealedCredential`) and the symmetric `src/primitives/index.ts` `seal()`/`unseal()`. diff --git a/02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md b/02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md new file mode 100644 index 0000000..4194c23 --- /dev/null +++ b/02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md @@ -0,0 +1,130 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-05 +deciders: jochen +reconstructed: false +--- + +# 49. A consumer's identity is bounded by the tightest backend that must accept it + +## Context + +The mesh says who a consumer is, once, and hands the same name to the provider (to create) and the +consumer (to present), so the two ends agree by construction rather than by two conventions (the +principle behind `ConsumerIdentity`, 04-ISSUES/023). The name is `mesh__`, cleaned to +lower-case letters, digits and underscore. + +Proving the provider contract per backend (ADR 0048) turned up 04-ISSUES/034: redis and postgres +create that name verbatim, but **minio refuses it** — an S3 access key is capped at 20 characters, +and `mesh_anchor_bucketuser` is 22. The provisioner then retries for ever, per consumer, and the +consumer holding that same too-long name could never present it either. + +Two things about the existing derivation decide most of this: + +- **The charset is already right.** `[^a-z0-9_]` is deliberately conservative, and its own comment + says it reaches "a PostgreSQL role, a MinIO access key, an LDAP uid and a Keycloak client without + quoting." That much is true. +- **The length is wrong.** `CheckIdentity` refuses names over `identityLimit = 63`, commented as + "the shortest identifier limit among the systems these names reach: PostgreSQL's". It is not the + shortest — S3's 20 is shorter — so the guard that was meant to catch exactly this lets it through, + and the failure lands at provision time as a silent retry instead of at assignment as a refusal. + +So this is a small wrong constant with a real cost attached: whatever bound we set, `mesh_` (5) plus +a node name plus `_` plus a module name has to fit inside it. + +## The options + +**A — Bound the identity by the true minimum, and refuse early.** Lower `identityLimit` to the real +shortest (20, S3's), so `CheckIdentity` refuses an over-long name *at assignment* with a clear +message, the way it already refuses over-63 names. The derivation does not change; long names are +simply rejected before anything is provisioned. +- *For:* smallest change; keeps "the mesh says the identity once, verbatim" intact; the failure + moves from a per-consumer provision-time retry to an up-front, legible refusal — which is what + `CheckIdentity` exists to do. +- *Against:* a hard budget. `mesh_` + node + `_` + module ≤ 20 means node + module ≤ 14 characters. + `anchor` + `bucketuser` (16) is already over. It pushes the constraint onto how machines and + modules are named, which is a real limitation on legible names. + +**B — Keep the readable name when it fits, compact it when it does not.** Below the bound, the name +is `mesh__` as today; over it, the mesh substitutes a deterministic short form (e.g. +`mesh_` + a truncated hash of node+module) — still one derivation, so both ends still agree. +- *For:* no naming constraint; short backends always satisfied; the common case stays legible. +- *Against:* some identities become opaque, and a provisioner tracing "whose login is this" loses + the answer for exactly the consumers that overflowed. The mesh now owns a fallback format and its + collision properties (a truncated hash is not free of collisions at 15 characters). + +**C — Let each interface declare its identifier bounds, and derive within the tightest a consumer +reaches.** `s3-bucket` states `identifier: { max: 20 }`; `postgres-database` states 63; the mesh +derives a name that fits the **minimum** bound across the providers a given consumer is granted. +- *For:* the most precise — each provision gets exactly the room it has, and a database consumer + keeps long legible names while an S3 consumer gets a short one; the constraint lives where the + fact does (on the interface). +- *Against:* the most work, and a consumer of two interfaces with different bounds must satisfy the + smaller — so its name shortens for both, reintroducing B's opacity in a narrower case. It also + means one consumer can hold **different** identities per provision, which the "said once" model + currently forbids. + +**D — Let the provider generate a backend-valid identity and hand it back (rejected).** minio mints +its own access key and returns it to the consumer. This is the data-provision return path this era +keeps meeting — but it directly contradicts 023 and ADR 0048: the identity would no longer be the +mesh's single derivation the two ends share, it would be a value one side invents and the other must +be told. Listed for completeness; not recommended. + +**E — A module (and a node) may declare a short slug; the identity is built from it.** The identity +becomes `mesh__`: where a slug is declared it is used, +otherwise the cleaned name. A slug is a deliberately short, operator-chosen identifier — `kc` for +keycloak, `wkstn` for a workstation. It is optional: short names (`anchor`, `redis`) need none. +- *For:* this is the escape hatch B wanted to be, without the opacity. The name stays legible — a + provisioner can read `mesh_wkstn_kc` and know who is asking — because a person chose it, not a + hash function. And it makes an early refusal *palatable*: if even the slug-built identity overflows, + the refusal points at the slug, a field made for exactly this, rather than at the machine's name. + Both ends still derive it from one declared thing, so they agree by construction. +- *Against:* a new optional manifest field, and someone must pick the slug — but only for names that + would otherwise overflow, and picking a short legible identifier is a better job than being handed + a hash. + +## What implementing A revealed + +A was tried first. At `identityLimit = 20`, the readable budget is `mesh_` (5) + node + `_` + module +≤ 20, i.e. **node + module ≤ 14 characters** — far tighter than it looked. The catalogue's own +existing tests use `workstation`+`keycloak` (25), which compacts to `mesh_dbbc02f8dde34d3`; common +mesh names (`home-server`, `the-build-node`, `workstation`) blow the budget with any module. So B's +compact fallback would fire for the *common* case, not the rare overflow — which inverts A+B: most +identities would be opaque hashes. A alone (hard refusal at 20) would refuse most realistic names. +This is what moved the recommendation to E: the problem is not the limit, it is that the *readable +name* is the wrong source when it is long, and a slug is a better source than either a hash or a ban. + +## Recommendation + +**E, over a per-consumer bound (start with the global minimum, 20).** Build the identity from an +optional slug, keep it when it fits, and refuse at assignment with "declare or shorten ``'s +slug" when it does not — no hash, no lost legibility, and the fix is a first-class field. Set the +bound to the true minimum (20) now; it needs no per-interface machinery to unblock S3, and a module +that consumes S3 simply declares a short slug. Graduate to **C** (per-interface bounds) later if it +turns out that non-S3 consumers are paying for S3's limit often enough to mind — E and C compose: +slugs are the mechanism, per-interface bounds refine where the ceiling sits. **B is dropped**: a +declared slug is a strictly better escape hatch than an opaque hash. **D stays rejected.** + +## Consequences (of E) + +- A module manifest gains an optional `slug`; a node may carry one too. `ConsumerIdentity` prefers + the slug over the cleaned name for each half. `identityLimit` becomes 20 (the true minimum), and + `CheckIdentity` refuses at `module add` / assignment — now with a message naming the slug to set. +- The common case stays legible; only names that overflow the budget need a slug, and what they get + is a name a person chose, not a hash. +- Existing modules/nodes whose names overflow declare a slug once — a migration cost paid as a clear + refusal with an obvious remedy, not a silent hash or a silent truncation. +- minio (04-ISSUES/034) is unblocked: an S3 consumer declares a short slug and its access key fits. +- **How it is checked:** the minio grant e2e — a consumer whose (slugged) identity fits reaches its + bucket with the credential the mesh delivered — plus unit tests that a slug is preferred, that an + un-sluggable over-long identity is refused (naming the slug), and that two consumers never collide. + +## References + +- [04-ISSUES/034](../04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md) — the + observation. +- ADR 0048 — a provider creates the credential the mesh minted; the identity it creates it under is + the one this decision bounds. +- `mesh-control` `internal/catalogue/identity.go` — `ConsumerIdentity`, `identityUnusable`, + `identityLimit`, `CheckIdentity` — where the constant and the check live. diff --git a/02-DECISIONS/0050-model-access-is-vendor-agnostic.md b/02-DECISIONS/0050-model-access-is-vendor-agnostic.md new file mode 100644 index 0000000..7952015 --- /dev/null +++ b/02-DECISIONS/0050-model-access-is-vendor-agnostic.md @@ -0,0 +1,206 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-05 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0024-model-access-is-a-provision.md +--- + +# 50. Model access is vendor-agnostic, and a vendor is an adapter + +## Context + +[ADR 0024](0024-model-access-is-a-provision.md) settled that model access is a provision and that +a licence is a named thing an operator uses. What shipped, and runs, is a single vendor: the mesh's +"claude" feature. A read-only trace of that feature (2026-09-05, in the code workspace) was made to +answer whether the model-access provision is Anthropic-shaped or genuinely general. The finding is +that **the vendor-agnostic layer already largely exists**, and the Anthropic specifics are a thin +band around it that a per-vendor adapter can hold. + +**What is already general, with evidence.** `mesh-control internal/licences` models +`licence(name, vendor, serves)` and `licence_holder(licence, node, module, sealed)`, and each +holder's credential is sealed per-holder through `internal/secrets`. `serves` carries the +non-secret facts (a base URL, a model) and is not vendor-specific. The `accept` verb +([ADR 0024](0024-model-access-is-a-provision.md), and +[`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)) already takes an operator-supplied +value, seals it to each holder and discards the plaintext. A model the mesh runs itself answers +`model-access` at node scope with no licence at all. None of that mentions Anthropic. + +**What is Anthropic-specific.** The credential is not a static key: it is a subscription OAuth grant +— an hourly access token plus a refresh token. That shape drags four things behind it that a static +key does not need: **central rotation** (one manager node refreshes under a lease and publishes the +new token), **delivery that strips the refresh token** so a consuming node holds only an access token, +an **identity guard** that reads the credential to catch a mis-binding, and a **usage** reading with +Anthropic's own `utilization%` semantics. Most vendors are a single static key, which the sealed-key +model already handles and which needs none of these four. + +**The tension at the centre of this.** A refreshable credential cannot be both *sealed so the mesh +cannot read it* and *rotated centrally*. Central rotation means some node in the mesh holds the +refresh token in readable form, because that is what refreshing requires. Per-holder sealing means no +node but the holder can read the credential. For a static key the two never meet — there is nothing to +rotate. For a refreshable grant they collide directly, and this record exists to say which gives way, +and by how much. + +## Considered Options + +1. **Keep Anthropic special-cased in the core.** Leave the three binding columns and the `claude_*` + schema, and add other vendors beside them the same way. **Rejected.** It is exactly what + [ADR 0024](0024-model-access-is-a-provision.md) ruled against: a module that names a vendor cannot + be moved onto another model without editing it, and moving it is the point. It also grows the core + by one band per vendor, when the bands are the same shape. + +2. **One provision, and refuse to hold any refresh token — re-seal only.** Make every credential + purely sealed per-holder, including refreshable ones; let each holder refresh its own grant. + **Rejected.** It throws away the hard half [ADR 0024](0024-model-access-is-a-provision.md) says + already works — the lease, the single-refresher, the switch-on-exhaustion — and replaces it with N + nodes each holding a refresh token, which is the very thing today's delivery strips on the stated + ground that *a node never holds a refresh token*. A refresh token is the long-lived secret; spraying + it across every holder is strictly worse than keeping one copy on one node. + +3. **One provision, and abandon central rotation entirely** for refreshable vendors — treat the grant + as opaque and let it expire. **Rejected.** For a subscription-seat vendor an expired access token is + a dead licence; without rotation the feature that works today stops working. This is option 2's cost + without option 2's autonomy. + +4. **One vendor-blind provision, plus a per-vendor adapter, with a bounded carve-out for the + refreshable case.** **Adopted**, below. + +## Decision + +**`model-access` is one consumer-facing, vendor-blind provision.** A consumer declares +`requires: model-access`, and is coupled to *reaching a model* — a base URL, a model name, a key — +and not to which vendor answers. That is the coupling the name is drawn at +([ADR 0040](0040-what-a-module-is.md)'s rule: name the interface at the widest boundary across which +the consumer does not care which implementation serves it). Where a consumer were genuinely coupled to +a specific wire API it could not swap across, the same rule would split the name — but the consumers +that exist reach their model through a CLI or SDK that hides the vendor, so `model-access` is the true +coupling and stays one name. This extends [ADR 0024](0024-model-access-is-a-provision.md) and +[ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) without changing them. + +**A vendor is an adapter, keyed by the licence's `vendor` field.** The lifecycle a licence needs is +vendor-specific and lives in a per-vendor adapter selected by `licence.vendor`, exactly as +`public-dns` is one neutral interface answered by registrar-scoped providers — +`cloudflare-dns`, `route53-dns` ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)). +A consumer names `model-access` and never a vendor, the same way a module names `public-dns` and never +a registrar. + +**The field is named `vendor`, not `provider`.** The inventory already uses "provider" for the +provider-pin — *which node answers a brokered provision*. Reusing it for *which company sells this +licence* would collide two unrelated facts on one word. `vendor` is the licence's, and is separate. + +### The adapter's capabilities, all but one optional + +An adapter declares: + +- **`shape`** — `static-key` or `refreshable-grant`. This is the switch the carve-out below turns on. +- **`accept(value) → sealed`** — take an operator-supplied credential and seal it to the holders, the + `accept` verb [ADR 0024](0024-model-access-is-a-provision.md) already defines. +- **`refresh(licence)`** — refreshable-grant only: the lease / rotate / publish machinery. +- **`identity(credential) → account-id`** — the mis-binding guard, for a vendor whose credential + carries an identity worth checking. +- **`usage(licence) → normalised rows`** — the vendor's usage reading, mapped to the common shape below. +- **`deliver`** — the credential *value* only; the destination path is the consumer's, not the + adapter's. + +**A static-key vendor implements almost nothing** — `shape: static-key`, `accept` is the generic +seal, `deliver` is the value, and `refresh`, `identity` and `usage` are absent or trivial. The +abstraction earns its keep by making the common vendor small, not the rare one clever. + +### The carve-out — the one place the guarantee is relaxed, said plainly + +The mesh's standing principle is that it cannot read what it stores: `accept` seals to the holders and +discards the plaintext ([ADR 0024](0024-model-access-is-a-provision.md)), and a provider seals nothing +because the credential travels the mesh's own asymmetric channel +([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)). A `refreshable-grant` +credential cannot honour that principle and be centrally rotated at the same time, and central rotation +is the working half [ADR 0024](0024-model-access-is-a-provision.md) is explicit about keeping. + +**So, for `refreshable-grant` vendors only:** + +- the **manager node holds the refresh token encrypted at rest** — readable by that node, because + rotation requires it. This is the bounded exception. +- **access tokens are still sealed per-holder**, as every credential is; a holder reads its own and no + other node reads it. +- the **refresh token is stripped on delivery** — it never reaches a consuming node. *A node never + holds a refresh token* stays true for every node but the one manager. + +**Static-key vendors keep the full guarantee.** There is no token to rotate, so there is nothing to +hold readably, so `accept` discards the plaintext and the carve-out never fires. The majority of +vendors are static-key, and the majority therefore lose nothing. + +The exception is stated rather than hidden because a relaxed guarantee that is not written down is +indistinguishable from a broken one. It is bounded on three axes at once: **refreshable-grant vendors +only, the refresh token only, the manager node only.** + +### The settled details this record also fixes + +- **Usage is normalised to `(licence, consumer, period, metric, value)` plus the raw response as + jsonb.** The metric is vendor-defined — Anthropic's `utilization%` is one metric, a token count is + another — and no common unit is forced across vendors. The raw response is kept so a reading can be + re-derived if the normalisation is later found wrong. +- **Binding is explicit per consumer, and an unchosen consumer is refused — no implicit fallback.** + This is the direction the resolver already takes, and it is the safe one: a mesh with several ways + to reach a model refuses a consumer that has not said which, naming the candidates and the command, + rather than silently choosing one ([ADR 0024](0024-model-access-is-a-provision.md), and + [`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)). +- **Subscription-seat authentication lives entirely inside the adapter**, never in the generic core. + So does an interactive `/login` — an adapter-specific "adopt" origin for a credential a person must + produce in a browser; the generic `licence key ` covers the static-key case. +- **Anthropic is the first `refreshable-grant` adapter**, carrying the OAuth refresh, the usage + reading, the identity guard and access-token-only delivery. **`anthropic-api-key` is a + `static-key` adapter for the same vendor's plain API keys**, and is the early second case that + proves the abstraction is not a single vendor wearing a coat: it exercises the whole path with the + carve-out switched off. + +## Consequences + +- **The carve-out is the mesh's one deliberate relaxation of "it cannot read what it stores."** It is + bounded to refreshable-grant vendors, to the refresh token, and to the manager node; the static-key + majority keep the full guarantee unchanged. This is the open risk the analysis carried here, and it + is recorded as an exception rather than pretended away. +- **Anthropic collapses from special case to adapter.** The three binding columns + (`nodes.node_license`, `nodes.hal_claude_account`, `agents.claude_account`) become three ordinary + consumers of `model-access`; the `claude_*` schema becomes the generic licence tables plus one + adapter. What was hardcoded becomes data keyed by `vendor`. +- **Adding a vendor is adding an adapter, and a static-key vendor is nearly free.** The modules that + want a model do not change when a vendor is added — they named `model-access`, not a vendor. +- **The refresh-token concentration is now a stated property to defend, not an accident.** The manager + node is a place a long-lived secret lives readably, and losing it or compromising it is a bounded, + named blast radius rather than a surprise. + +### How each claim here is checked + +- **Vendor-blind provision, static-key path.** A lab scenario binds an `anthropic-api-key` licence to + a consumer; the consumer resolves, receives a key sealed to its node, and reaches a model — and the + key is **nowhere in the control plane's database** nor in anything that crossed the broker. This is + the `licence_holder` sealed-per-node check that [`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md) + already runs, now asserted for a second vendor. +- **The carve-out is exactly as narrow as stated.** For a `refreshable-grant` licence, a test asserts + the refresh token exists (encrypted) **only on the manager node**, is **absent from every holder's + delivery**, and that the delivered credential is access-token-only — and that for a `static-key` + licence no refresh token is stored anywhere. +- **Adapter selection is keyed by `vendor`.** A scenario with two vendors on two licences verifies each + licence's lifecycle runs its own adapter, and that a consumer naming `model-access` never names a + vendor to get one. +- **Refuse-if-unchosen.** Already checked in [`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md): + a consumer with more than one candidate is refused with the candidates and the command named. + +## References + +- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision, a licence is a named + thing, and `accept`; this record generalises its single vendor and keeps its working central + rotation. +- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — a provision names the + coupling; `model-access` is drawn at the consumer's. +- [ADR 0040](0040-what-a-module-is.md) — the naming rule and the neutral-interface / scoped-provider + shape a vendor adapter follows. +- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — registrar-scoped `public-dns` + providers, the precedent a `vendor`-scoped adapter mirrors. +- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md) — the mesh seals credentials + and holds no readable copy; the carve-out here is the bounded, named exception to that for a + refreshable grant. +- [`03-DESIGN/01-to-be/14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) — the design this + record extends, amended to describe the adapter generalisation. +- The read-only vendor-agnostic analysis, 2026-09-05 (code workspace) — the inventory and the decisions + taken on the open questions this record encodes. diff --git a/02-DECISIONS/0051-shared-data-is-the-operators.md b/02-DECISIONS/0051-shared-data-is-the-operators.md new file mode 100644 index 0000000..ee7b9ce --- /dev/null +++ b/02-DECISIONS/0051-shared-data-is-the-operators.md @@ -0,0 +1,166 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-05 +deciders: jochen +reconstructed: false +extends: 0030-data-outlives-the-mesh-that-declared-it.md +--- + +# 51. Shared data is the operator's, and a module is granted access to it + +## Context + +**Eight modules declared one filesystem as eight private ones.** The media stack — a library +server, the acquisition managers for films, series, music and books, a subtitle fetcher and two +download clients — shares directories on one machine: the download clients write into +`/services/media/downloads` and the managers read it; the managers write into the libraries and +the library server reads them. That sharing is the entire point of the stack. Yet each module +declared every shared directory it touched as its own `directory` resource, with an owner and a +mode. `/services/media/downloads` was written seven times, as seven private directories that +happen to be the same path. + +**The resolver refuses exactly that, and is right to.** Two modules declaring one path on one node +are refused by name, with no exemption for identical content and no merge — because two owners of +one path is the class of fault this repository keeps recording ([04-ISSUES/036](../04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md)). +So the stack as written refuses its own only sensible assignment: all of it on one machine, +sharing one filesystem. It passed today only because no test co-resolves any two of the eight. The +first machine assigned two of them together is where the refusal would have surfaced. + +**The vocabulary had one word for two intentions, and this was already seen.** +[04-ISSUES/026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md) found the +same gap from the other side and named it precisely: two kinds of mount are spelled identically — +*the directory my data lives in*, which the mesh creates and owns, and *a facility I was granted*, +which already exists and the mesh only reaches. That issue deferred inventing a field to tell them +apart, because doing so is a design decision and it declined to make one to get a check green. This +is that decision. + +**A `directory` resource is owned, on every axis.** The host creates it, sets its owner and mode, +and removes it when it is empty and no longer declared — it *removes what it made and leaves what +it merely configured* ([ADR 0005](0005-the-node-host.md), [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)). +A shared media library is none of that. It existed before the mesh, several modules read and write +it at once, and losing it is the one failure that does not recover. It is not any module's +resource; it is the operator's, and a module only needs to be let at it. + +## Considered Options + +1. **Make the stack one module with several containers**, the way the mail system already is. The + issue raises it directly: is a set of modules that must share a filesystem really one module? + **Rejected.** The eight are independently assignable and independently useful — a person may run + the download client without the library server, or the film manager without the music one — and + folding them into a single module to express a shared directory would make *what a module is* + turn on an incidental filesystem contract. It also does not generalise: the next pipeline of + modules handing files to each other on one machine (an ingest folder, a spool, a drop directory) + would face the same wall and the same wrong remedy. + +2. **One module owns the directories and the rest `require` them.** **Rejected**, and this is the + heart of the decision. Nobody owns shared operator data. The library predates the mesh and + outlives any one module, so making the library server or a manager its owner means unassigning + that module orphans everyone else's access — and the owner would set the owner and mode of a + tree it did not create. Ownership is the wrong relationship to model, because the true owner is + not a module at all. + +3. **A flag on a directory resource** — `external: true`, or an owner of `operator`. **Rejected.** + It overloads one shape with a boolean that inverts every one of its semantics: created becomes + *must already exist*, owned becomes *touch nothing*, removed-when-empty becomes *never removed*. + That is the *two-kinds-spelled-identically* trap of 04-ISSUES/026 reintroduced with a single + quiet field — a reviewer reading `type: directory` would have to check one boolean elsewhere to + know whether the mesh owns the thing at all. + +4. **A distinct `accesses` declaration, separate from resources.** **Adopted.** + +## Decision + +**Shared, pre-existing data is operator-owned and external. The mesh does not create it, does not +set its owner or mode, does not reconcile it and does not remove it.** A media library, a download +spool, an ingest directory is the operator's, and the mesh is a guest in it. + +**A module declares that it needs *access* to such a path, not that it owns a resource there.** The +manifest field is `accesses`: a list of `{path, mode}`, where mode is `read` or `read-write` and +absent narrows to `read` — the safe default, because the danger with an access is being given more +than was meant, not less. A module's own configuration and state directories stay owned +`directory` resources; only the shared, pre-existing paths become accesses. + +**The host mounts an accessed path and owns nothing about it.** It reaches the machine as a new +declaration shape, `access`, distinct from `directory`. The host confirms the path is present and +does nothing else — no create, no chown, no mode, no removal. + +**An accessed path absent at apply time is refused, clearly, not created.** The mesh does not own +it, so conjuring it would be a lie the host then acts on — and specifically the lie 04-ISSUES/026 +records, where a bind mount whose source does not exist is made by the container runtime as root +with the wrong ownership. The host says the operator must provide the path instead. + +**Several modules accessing one path is normal, and never refused.** The duplicate-path refusal is +about *ownership*, not *use*: it applies to resources a module owns and to those alone. An access +is not a resource and never enters the check, so the eight-module stack co-resolves. What stays +refused is genuine rivalry — two modules owning one path — and the new contradiction it exposes: a +path one module owns while another merely accesses it, because that asserts both that the mesh owns +the directory and that the operator does. + +This is a decision and not a patch because it settles *what a module may say about a path it did +not make*, which every co-located file-handoff in the catalogue now and later depends on — and +because it draws the ownership line [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) +started: the host owns what it made and keeps what it merely configured, and this adds the third +case it did not have a word for — what it neither made nor configured, and must not touch. + +### How each claim is checked + +- **The stack co-resolves.** A control-plane unit test assigns two modules that declare access to + one path on one node and asserts no refusal — the exact case the resolver refuses when the same + path is owned. The mirror test, two modules *owning* one path, still refuses, so the sharing + vocabulary does not weaken the rule it sits beside. +- **Ownership and access cannot both be claimed of one path.** A unit test asserts the resolver + refuses a path one module owns and another accesses, naming both. +- **Absent is refused, not created.** A host unit test applies an access to a path that does not + exist and asserts a clear refusal that names the operator, and that nothing was created. +- **Present is confirmed and nothing moves.** A host unit test applies an access to an existing + directory and asserts the apply reports no change and disturbs nothing. +- **Undeclaring never removes.** A host unit test drops a previously declared access and asserts + the operator's directory and its contents are left exactly as they were — the data-loss failure + [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) exists to prevent, on a directory the + mesh never made. +- **Every media manifest is corrected.** No `/services/media/*` path is an owned `directory` + resource in any of the eight; each is an `accesses` entry, and each module's own config and state + directories remain owned. Checked by the control-plane manifest parser, which now understands + `accesses` and refuses a malformed one. + +## Consequences + +**A shared filesystem between co-located modules now has a vocabulary**, and it is not the media +stack's alone: any pipeline handing files to a neighbour on one machine — an ingest directory, a +spool, a drop folder — says *I access this operator path* rather than *I own this directory*, and +several of them may say it of one path. + +**Unassigning a module that reached shared data leaves the data.** Correct, and the same trade +[ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) made for owned directories: removing +data is a person's act, done knowingly, not a side effect of unassignment. + +**The operator must provision the shared paths before the stack is applied**, and a machine that +lacks one is told plainly which. That is a real new obligation, and it is the right one: the mesh +cannot own what predates it, so it cannot create it either, and saying so at apply time beats a +directory conjured as root and a service that half-works. + +**The host vocabulary grew by one shape**, which is a cost — every added shape widens what a +compromised control plane can express ([ADR 0005](0005-the-node-host.md)). It is a narrow one: an +`access` is confirmed by a stat and grants the host no new action. It earns its place by letting +the host refuse to create what it must not own, which no existing shape could say. + +**A path can be both owned and accessed only by refusal.** If a future manifest declares one path +as an owned directory in one module and an access in another, the resolver refuses it rather than +guessing which is meant — the two assertions about who owns the data cannot both hold. + +## References + +- [04-ISSUES/036](../04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md) — six (in + fact eight) modules own what they must share; the problem this resolves +- [04-ISSUES/026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md) — the two + kinds of mount spelled identically, which deferred this field to a decision +- [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) — data outlives the mesh; the host + keeps what it did not make. This record adds the case it had no word for +- [ADR 0005](0005-the-node-host.md) — the host removes what it made and leaves what it merely + configured; the vocabulary is finite and every shape is a security decision +- [ADR 0040](0040-what-a-module-is.md) — what a module is; an access is a new thing a module may + say about the machine it lands on +- mesh-control `feat/shared-data-access`, mesh-catalog `feat/media-access-not-ownership`, + mesh-host `feat/mount-operator-owned` — the mechanism, the corrected manifests, and the host + shape diff --git a/02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md b/02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md new file mode 100644 index 0000000..1f2e386 --- /dev/null +++ b/02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md @@ -0,0 +1,198 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-05 +deciders: jochen +reconstructed: false +extends: 0005-the-node-host.md +--- + +# 52. An init step is a container run once to completion, gating what follows + +## Context + +**A module can declare things that exist; it cannot declare a step that runs.** The host owns a +finite vocabulary of shapes — `file`, `directory`, `service`, `package`, `container`, `action` — +and every one but `action` describes *state*: a thing that should be present, with content or a +mode or an image, which the host reconciles toward ([ADR 0005](0005-the-node-host.md)). That is +right for what it covers. But a real class of modules needs, once, to *run their own code at a +point in their own lifecycle* — and the vocabulary has no word for it +([04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md)). + +**mosquitto is the sharp case, and it fails silently without this.** Its Dynamic Security plugin +loads at broker start and refuses to come up unless `dynamic-security.json` already holds an admin +client. Seeding that file is a step that must happen *after* the data directory exists and *before* +the broker container starts. The manifest can declare the directory, the config file and the broker +container; it cannot declare "seed this, once, before that container starts." Written as it is +today, the broker starts against an unseeded store and the plugin aborts — and the next reconcile +does not fix it, because nothing in the declaration ever seeds the file. + +**It is not one module's defect.** The database providers need the same to run a first-boot +migration, an extension enable, or a health gate before they are announced ready; today that works +only where the *image* happens to seed itself from an environment variable, and anything the mesh +must run once against the server has no home. This is the timing face of the same gap +[04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md) records +from the content side: the manifest needed *the file to exist before first start*, and had only +*the file has this content, forever*. + +**The obvious answer is the one that already went wrong.** An earlier mesh had exactly this as a +feature — event-driven hooks that ran custom code at phases of build, publish and deploy. It was +powerful and it was *complex to set up and flaky*, and that fragility, not the need, is the content +of the issue. Whatever this becomes must not rebuild that engine. + +**The ground has shifted since that engine, in a way that makes a much smaller answer possible.** A +module with tools or events now runs a **process of its own** — a container carrying the module's +compiled code, holding the single broker account the mesh scoped to it, isolated from every other +module ([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The +code that must run at first boot is *already in that image, under that account*. So the mesh does +not need a way to run a module's code — it has one. It needs a way to say **run this container to +completion, and start the one that depends on it only after it has.** + +## Considered Options + +1. **A host `run` shape — a command the host executes on the machine.** The direct reading of + "run code at a lifecycle phase." **Rejected.** The host's vocabulary is finite and every added + shape is a security decision, because it widens what a *compromised control plane* can express + ([ADR 0005](0005-the-node-host.md)). A general "run this command as the host" is the largest + such widening there is: the blast radius is the whole machine, as root. The mesh already drew + this exact line for `action` — it runs a command, and so it is *permitted from the bundle and + refused from the link*, because the bundle arrives with the binary and the link is a separate + party with an unbounded reach. A new host-command shape usable by an ordinary module would be an + `action` from the link by another name, which is precisely what is refused. + +2. **Per-phase lifecycle hooks on a module** — `pre-start`, `post-start`, `pre-remove`, and their + build/publish cousins, each naming code the mesh runs at that phase. The general answer, and the + old feature. **Rejected for now.** It is the flaky engine the issue warns against, and most of + its phases have no present need. Deciding the full set of phases, where each one's code runs, and + how each is made idempotent is a large design taken to buy capability nothing yet asks for. The + three blocked modules all need one phase — *before a container starts* — and a mechanism narrow + enough to be obviously correct beats a general one that is not. + +3. **A distinct one-shot resource type** — a new shape, sibling to `container`, that names an image + and runs it once. **Rejected.** It grows the host vocabulary by a whole shape (a `Type`, a + struct, an applier, a place in every host's shape list) to express something a `container` almost + already is. A one-shot *is* a container — a pinned image, an account, volumes, an environment — + that happens to exit. Spending a new shape on the difference is the cost of option 1 in smaller + type, for a capability the existing shape can carry with one modifier. + +4. **A modifier on the existing `container` shape: this container runs once, to completion, and the + host gates the apply on it.** **Adopted.** It reuses the shape the host already has, adds no new + host action, and leans on two guarantees the host already gives — *apply in declared order, + never sorted*, and *a failed step fails the apply* — to turn "before that container starts" into + an emergent property of ordering rather than a dependency graph the host must resolve. + +## Decision + +**A run-once step is an ordinary `container`, marked to run to completion.** The manifest sets +`run-once: true` on a container resource. Everything else about it is a container as before — a +digest-pinned image ([ADR 0006](0006-the-substrate-and-the-control-plane.md)), volumes, an +environment, and for a module's own code the same scoped account its runtime already holds +([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The mesh adds +no way to run code; it marks a container the host must **run to completion and require to exit 0**, +rather than start and leave running. + +**The gate is declaration order, not a named dependency.** The host applies a declaration in the +order it is given, does not sort, and does not resolve dependencies — ordering is a decision, and it +is the control plane's ([ADR 0005](0005-the-node-host.md), +[04-ISSUES/013](../04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md)). A +run-once step is placed *before* the container that depends on it, and **a failed run-once step +halts the apply**, exactly as a failed `action` does — so everything the declaration places after +it, the broker included, is never reached until the step has completed. "Before the broker starts" +is therefore expressed by list position plus completion, and the host cross-references nothing. + +**Completion is recorded, and a re-apply does not re-run it.** The host records what it applied only +after the fact, as the digest of the declaration that produced it +([ADR 0018](0018-a-picture-is-read-from-what-runs.md)) — a run-once step no differently. Because the +step leaves nothing running to inspect, that persisted digest, not a live container, is the marker +that it happened. On a later apply the host finds the digest already recorded for this exact +declaration and does nothing; it re-runs only when the declaration's digest has changed, and a step +that exited non-zero recorded nothing and so is retried next apply. This is the reconcilable, +idempotent discipline the state shapes get for free, made explicit for a step. + +This is a decision and not a patch because it settles **what a module may say about running its own +code**, which the whole catalogue of providers — a seed before start, a first-boot migration, a +health gate — now and later depends on, and because it draws the line the issue asked for: the +narrowest sound mechanism that unblocks the three modules without rebuilding the hook engine whose +fragility is the warning. + +### The security bound, stated plainly + +**The host gains no new action and no new shape.** `run-once` is a boolean modifier on the +`container` shape that already exists. A run-once container is strictly *less* powerful than an +`action`: it cannot run an arbitrary host command, only a digest-pinned image under an account the +mesh scoped — which is exactly the capability `container` already grants from the link. A control +plane that is compromised can express nothing through `run-once` it could not already express by +declaring an ordinary `container`. The dangerous expansion of option 1 — a command the host runs on +the machine — is not made. + +### How each claim is checked + +- **A run-once step runs to completion and its exit 0 is required.** A host unit test applies a + run-once container whose image exits 0, asserts the host ran it to completion (not detached, not + left running) and reported it done; a sibling test applies one that exits non-zero and asserts + the apply fails, naming the step. +- **A failed run-once step gates what follows.** A host unit test places a run-once container that + exits non-zero before another container and asserts the second is never started and the apply is + reported gated — the mirror of the existing test that a failed action stops what follows. +- **It is not re-run once it has completed.** A host unit test applies a run-once step, then applies + the identical declaration again with the first run's record present, and asserts the second apply + runs nothing and reports the step unchanged. +- **A changed declaration re-runs it.** A host unit test applies a run-once step, then applies one + whose image or environment differs, and asserts it runs again — the digest moved, so the marker no + longer matches. +- **The vocabulary carries the field end to end.** A control-plane unit test resolves a module whose + manifest marks a container `run-once` and asserts the rendered host declaration carries the field, + in author order before the container it gates; the manifest parser refuses a `run-once` that is + not boolean. +- **mosquitto seeds before the broker.** mosquitto's manifest declares a run-once init container, + before the `server` (broker) container, that writes the admin client into `dynamic-security.json` + and exits — checked by the control-plane resolver, which now understands the field, and by the + ordering of the rendered declaration. The end-to-end proof that it runs exactly once, at the right + phase, and converges on re-apply is owed to a lab scenario ([04-ISSUES/037] open question), which + this record does not close. + +## Consequences + +- **The three blocked modules gain a home for their step.** mosquitto seeds its dynsec admin before + the broker; a provider that must migrate or health-gate at first boot declares a run-once step in + its own runtime image, under its own account, before the container that depends on it. +- **The general lifecycle hook is deferred, deliberately.** Only *before a container starts* is + bought here. `post-start`, `pre-remove` and the build/publish phases remain unbuilt, and the day + one is genuinely needed it is decided then, against a need, not speculatively — the same restraint + that kept this from being the old engine. +- **The seed-then-mutate file is safe if the step is written to be.** A run-once seed writes + `dynamic-security.json` only when it is absent and never reconciles it, so what the running plugin + grows in that file afterward is never wiped + ([04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)). + The host's marker guarantees the step is not re-run; the step's own code guarantees it does not + clobber on the pass it does run. +- **The host vocabulary did not grow, and that is the point.** The cost of a run-once step is one + boolean and a completion path in the container applier, not a new shape and not a new action. The + mesh expresses ordering and completion; the module runs its own code, where it already runs it. +- **A run-once step that never converges is a stuck apply, loudly.** A step that exits non-zero + every time halts the apply every time, and the container it gates never starts — which is the + correct failure, reported, rather than a broker that half-starts against an unseeded store and a + reconcile that reports success. It is failed forward, not failed silent. + +## References + +- [04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) — a + module cannot run its own code at a lifecycle phase; the gap this resolves, and the warning about + the old hook engine +- [04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md) — the + content face of the same gap: a file needed before first start, that the running program then + mutates +- [04-ISSUES/013](../04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md) — the + order of a declaration is the control plane's, and the host applies it as given; the gate rests on + this +- [ADR 0005](0005-the-node-host.md) — the host's finite vocabulary, the ordered declaration it does + not sort, and `action` as the shape a command already is and why it is refused from the link +- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — a module runs + its code as its own process under its own account; a run-once step is that process, run to + completion +- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — what was applied is recorded after it works; + the completion marker is that record +- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — a container is pinned by digest; a + run-once container no differently +- mesh-control `feat/lifecycle-run-once`, mesh-host `feat/apply-run-once`, mesh-catalog + `feat/mosquitto-bootstrap` — the vocabulary, the apply support, and mosquitto's seeded broker diff --git a/02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md b/02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md new file mode 100644 index 0000000..b13d6e6 --- /dev/null +++ b/02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md @@ -0,0 +1,181 @@ +--- +topic: what runs on it +status: accepted +date: 2026-09-06 +deciders: jochen +reconstructed: false +extends: 0052-a-step-that-runs-once-before-a-container.md +--- + +# 53. A scheduled step is a container run on a recurring schedule + +## Context + +**[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) gave the mesh a step that runs *once*; +a real class of modules needs one that runs *again and again*.** kometa reconciles a media library +against its lists on a timer; a ticketing integration polls its source for new work every few minutes; +a backup, a cache warm, a metrics roll-up all recur. The host's vocabulary describes *state* — a file, +a directory, a container that should be running — and 0052 added *a step that happens once and is +done*. Neither says *this should happen every night at 3, forever* +([04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) named +the lifecycle gap; 0052 closed the once-before-start face of it and left the recurring face open). + +**The modules that need it already run their own code.** As with run-once, the ground has shifted +since the old mesh's flaky hooks: a module with tools or events runs a **process of its own** — a +container carrying the module's compiled code under the single scoped account the mesh gave it +([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The work that +must recur is already in that image, under that account. The mesh does not need a way to run a +module's code on a timer — it has the code and the account. It needs a way to say **run this +container again on this cadence**. + +**The shape of the answer is already decided, one modifier over.** 0052 rejected a host `run` command, +per-phase hooks, and a distinct one-shot resource type, and adopted *a modifier on the `container` +shape the host already has*, because a one-shot is a container that happens to exit. A scheduled step +is the same container that happens to exit — run repeatedly. Reusing that shape keeps the host +vocabulary flat and inherits 0052's security bound whole. The only thing 0052's `run-once` does not +carry is *when to run it again*. + +**One difference from run-once changes a rule, and it is the reason this is its own record.** A +run-once step **gates the apply**: it is placed before the container that depends on it, and a failure +halts everything after it, because "seed the store before the broker starts" is a correctness +precondition ([ADR 0052](0052-a-step-that-runs-once-before-a-container.md)). A scheduled step is the +opposite: it runs *after* the machine is up and converged, on its own clock, and a single failed run +is an ordinary operational event — the next run comes anyway. A scheduled step that halted the apply, +or that a failed run marked the node not-current over, would make a routine poll into a reason the +whole machine reads as broken. So the gating rule 0052 established is exactly the rule this record +must **not** inherit. + +## Considered Options + +1. **A host `cron`/`timer` shape — the host installs a system timer that runs a command.** The direct + reading. **Rejected**, for the reason 0052 rejected a host `run` shape: it widens what a + *compromised control plane* can express toward "run this command on the machine, forever," which is + the largest widening there is, and it is `action`-from-the-link by another name + ([ADR 0005](0005-the-node-host.md)). A recurring command is worse than a one-off, because it + persists. + +2. **Per-phase lifecycle hooks** — `on-schedule` joining `pre-start`/`post-start` as named code the + mesh runs. **Rejected for now**, as in 0052: it is the flaky hook engine the issue warns against, + and the three modules that need this need one thing — *run this container on a cadence* — which a + narrow modifier expresses without deciding a whole hook vocabulary. + +3. **A distinct `scheduled` resource type**, sibling to `container`. **Rejected**, as 0052 rejected a + distinct one-shot type: it spends a whole new host shape (a `Type`, a struct, an applier, a place + in every host's shape list) on something a `container` already almost is — a scheduled task *is* a + container (pinned image, account, volumes, environment) that runs on a clock. + +4. **A modifier on the existing `container` shape: `schedule`, a cron expression the host runs the + container on.** **Adopted.** It reuses the shape the host has, adds no new host action, and sits + beside `run-once` as its recurring twin — the same container, exited, run again. + +## Decision + +**A scheduled step is an ordinary `container`, marked with a `schedule`.** The manifest sets +`schedule: ""` on a container resource — a standard five-field cron expression. Everything else +about it is a container as before: a digest-pinned image +([ADR 0006](0006-the-substrate-and-the-control-plane.md)), volumes, an environment, and for a module's +own code the same scoped account its runtime already holds +([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The mesh adds no +way to run code; it marks a container the host must **run on that cadence, each time to completion**, +rather than start once and leave running (a service) or run once and gate (a run-once step). + +**A run is fired by the clock, not by the apply, and does not gate it.** Applying the declaration +installs the schedule; it does not run the step. The machine converges — reports applied and current — +as soon as the schedule is installed, exactly as it does for a service that is running. Thereafter the +host fires the container when the cron expression is due. This is the deliberate inversion of 0052: +a scheduled step is downstream of convergence, not a precondition of it. + +**A failed run is recorded and the next run still comes; it never marks the node not-current.** A run +that exits non-zero is logged against the module — visible, auditable — but it does not fail the apply, +does not halt other resources, and does not flip the node's reported state. A poll that fails at 03:00 +and succeeds at 03:05 is the system working, not a machine that needs attention. A step that fails +*every* time is a loud, repeating log entry, which is the correct signal for "this recurring job is +broken" — distinct from "this machine did not converge." + +**Runs do not stack.** If a run is still going when the next is due, the host skips the due run rather +than starting a second copy, and logs the skip. A slow nightly job that occasionally overruns must not +spawn a growing pile of concurrent containers competing for the same account and volumes — the failure +mode that made the old timers dangerous. + +**Each run is independent and idempotent by the module's own code.** The mesh guarantees only *the +container is run on the cadence*; that a run does the right thing when the previous one half-finished +is the module's contract, the same discipline a run-once seed owes +([ADR 0052](0052-a-step-that-runs-once-before-a-container.md), +[04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)). + +This is a decision and not a patch because it settles **what a module may say about running its own +code on a cadence**, which every recurring provider job — a sync, a poll, a roll-up — now and later +depends on, and because it draws the line 0052 left: the recurring twin of run-once, with the gating +rule deliberately reversed so a routine job's failure is never a machine's failure. + +### The security bound, stated plainly + +**The host gains no new action and no new shape.** `schedule` is a string modifier on the `container` +shape that already exists. A scheduled container is strictly *less* powerful than an `action`: it runs +a digest-pinned image under an account the mesh scoped, which is exactly what `container` already +grants from the link, and it cannot run an arbitrary host command. A compromised control plane can +express nothing through `schedule` it could not already express by declaring a `container` — the cron +string only says *how often*, not *what*. The dangerous expansion of option 1 — a command the host +runs on the machine on a timer — is not made. + +### How each claim is checked + +- **A scheduled step runs when the schedule is due.** A host unit test installs a container with a + schedule that is due immediately (or advances a injected clock to when it is due) and asserts the + host ran it to completion; a sibling test with a schedule not yet due asserts it has not run. +- **Installing it does not run it, and the node is current without a run.** A host unit test applies a + scheduled container and asserts the apply reports current *before* any run has fired — the schedule + is state that is present, not a step that gated. +- **A failed run does not fail the apply or the node.** A host unit test fires a scheduled container + that exits non-zero and asserts the failure is recorded against the module, the apply is not failed, + and the node stays current — the mirror of the run-once test where a non-zero exit *does* halt. +- **Runs do not stack.** A host unit test fires a scheduled container whose run outlasts its next due + time and asserts the host skipped the due run and logged the skip, rather than starting a second + container. +- **The vocabulary carries the field end to end.** A control-plane unit test resolves a module whose + manifest sets `schedule` on a container and asserts the rendered host declaration carries the field; + the manifest parser refuses a `schedule` that is not a valid cron expression, and refuses a container + that is both `run-once` and `schedule` (a step is one or the other, never both). +- **A real module recurs in the lab.** A converted module declaring a scheduled step (kometa's library + sync, or a poller) is assigned in a lab scenario, and the scenario asserts the scheduled container + fires on its cadence and its account and volumes are the module's — the end-to-end proof this record + owes, as 0052 owed its run-once lab proof. + +## Consequences + +- **The recurring providers gain a home for their cadence.** kometa reconciles on its schedule; a + poller polls; a roll-up rolls up — each a scheduled step in its own runtime image, under its own + account, on the cron it declares. +- **The general lifecycle hook is still deferred.** Only *run once before* (0052) and *run on a + cadence* (this) are bought. `post-start`, `pre-remove` and the build/publish phases remain unbuilt, + decided when a real need arrives, not speculatively — the restraint that kept both from being the old + engine. +- **A recurring job's failure is loud but not fatal.** The node stays current while a scheduled step + fails and retries; a step that fails forever is a repeating log entry, not a machine marked broken. + This is the correct separation — a machine's convergence and a job's success are different questions — + and it is why this could not simply be `run-once` without the schedule. +- **The host vocabulary did not grow, again, and that is the point.** The cost of a scheduled step is + one string field and a cron loop in the container applier, not a new shape and not a new action. The + mesh expresses cadence; the module runs its own code, where it already runs it. +- **`run-once` and `schedule` are exclusive and complete for now.** A container runs once and gates, or + runs on a cadence and does not, or runs and stays up (a service). A manifest that asks for two of + these at once is refused, because the three are distinct answers to "how does this container run." + +## References + +- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the run-once step; this is its + recurring twin, reusing the `container`-modifier shape and inheriting its security bound, and + reversing its gating rule +- [04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) — a + module cannot run its own code at a lifecycle phase; 0052 closed the once-before-start face, this + closes the recurring face +- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — a module runs + its code as its own process under its own account; a scheduled step is that process, run on a cadence +- [ADR 0005](0005-the-node-host.md) — the host's finite vocabulary, and why a host command (even on a + timer) is refused from the link +- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — a container is pinned by digest; a + scheduled container no differently +- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — what was applied is recorded after it works; + the installed schedule is state, a fired run is an event +- mesh-control `feat/schedule-container`, mesh-host `feat/apply-schedule`, and the converted module + that first declares a scheduled step — the vocabulary, the apply support, and the recurring proof diff --git a/02-DECISIONS/0054-model-usage-is-recorded-at-two-grains.md b/02-DECISIONS/0054-model-usage-is-recorded-at-two-grains.md new file mode 100644 index 0000000..5e17331 --- /dev/null +++ b/02-DECISIONS/0054-model-usage-is-recorded-at-two-grains.md @@ -0,0 +1,178 @@ +--- +topic: model access +status: accepted +date: 2026-09-06 +deciders: jochen +reconstructed: false +extends: 0050-model-access-is-vendor-agnostic.md +--- + +# 54. Model usage is a vendor-neutral record, produced by the adapter, at two grains + +## Context + +**[ADR 0050](0050-model-access-is-vendor-agnostic.md) fixed the *shape* of a usage reading and left +its *home* open.** It decided that an adapter may expose `usage(licence) → normalised rows`, that a +row is `(licence, consumer, period, metric, value)` plus the raw vendor response as jsonb, and that the +metric is vendor-defined (Anthropic's `utilization%` is one metric, a token count another). It did not +say where those rows are stored, how they get produced on a cadence, or whether the only grain is the +whole licence — and the mesh's first real consumer needs all three answered. + +**The predecessor mesh recorded usage at two grains, and both are wanted.** It polled the vendor for a +licence-level reading (an account's `utilization%`), and it also attributed **per-session** token and +cost — which model, how many input and output tokens, what it cost — parsed from the agent's own +transcript, including the case where a long session switched the account it billed against mid-way. The +licence-level reading answers "how close is this subscription to its cap"; the session-level reading +answers "what did this piece of work cost, and against which account." A mesh that kept only the first +could not bill a project or notice a runaway session; keeping only the second could not see a cap +approaching. Both are load-bearing and neither subsumes the other. + +**A session is already a consumer, so the second grain needs no second vocabulary.** +[ADR 0026](0026-the-mesh-has-a-session-of-its-own.md) and design 15 establish that an agent session is +a thing in its own right, and that **model access is bound to the session, not to the machine** — a +session is a *consumer* of `model-access`, identified by `(node, module)` and its own id. That is +exactly the `consumer` column ADR 0050 already put in the usage row. So the two grains are not two +schemas; they are the same row at two consumer resolutions: the holding module for the licence grain, +the session for the finer one. The design's own still-open worker-naming gap (design 14) is the same +gap here and is left where it is — a session id distinguishes what `(node, module)` cannot. + +**The pieces this needs already exist.** A periodic reading is a **scheduled step** +([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)) — the adapter's `usage()` poll is a container the +mesh runs on a cadence, which is precisely what 0053 was built for. A durable audit of what happened is +**an event the audit trail records** ([ADR 0041](0041-events-are-a-relationship.md), +[ADR 0042](0042-the-shape-of-an-event-on-the-wire.md)); the `audit-logger` module already consumes every +event. And a queryable current picture is **a context store**, the same shape the licences themselves +live in ([ADR 0008](0008-a-context-owns-its-store.md)). Nothing new in kind is required; what is missing is +the decision to point them at usage. + +## Considered Options + +1. **Licence grain only — a single `utilization%` poll, nothing per session.** Rejected: it cannot + attribute cost to a piece of work or catch a session that is burning an account down, which is half + of why usage is recorded at all. + +2. **A bespoke `sessions` schema mirroring the old mesh's `*_sessions` / `*_session_account_usage` + tables.** Rejected: it reintroduces a second, vendor-shaped vocabulary for something the mesh + already names — a session is a consumer, and its usage is a usage row. A parallel schema would drift + from the `model-access` vocabulary and force every reader to learn two. + +3. **Store usage only as raw vendor blobs, normalise later.** Rejected as the *whole* answer (kept as a + fallback within the chosen one): a reader that must parse Anthropic's response shape to answer "what + did this cost" has the vendor coupling the whole feature exists to remove. The raw blob is kept + beside the normalised row (0050 already requires this), not instead of it. + +4. **One vendor-neutral usage record at two consumer grains, produced by the adapter, recorded as + both an event and a queryable row.** Adopted. + +## Decision + +**Model usage is one vendor-neutral record — `(licence, consumer, period, metric, value)` plus the raw +response — recorded at two grains that differ only in the `consumer`.** At the **licence grain** the +consumer is the holding module and the metric is the vendor's own account reading (Anthropic: +`utilization%`). At the **session grain** the consumer is the agent session — `(node, module)` and its +session id ([ADR 0026](0026-the-mesh-has-a-session-of-its-own.md)) — and the metrics are the ones a +session bills: input tokens, output tokens, model, and cost. The row shape is 0050's, unchanged; the +grain is which consumer the row is *for*. + +**The adapter is the only thing that knows the vendor, and it produces both grains.** Reading an +account's cap is the adapter's `usage(licence)` verb ([ADR 0050](0050-model-access-is-vendor-agnostic.md)); +attributing a session's cost is the adapter reading that vendor's transcript or usage API and emitting +rows keyed to the session. The mesh defines the row and the plumbing; the adapter fills it from whatever +the vendor exposes, and a static-key vendor that exposes nothing simply produces no rows — usage is an +optional reading, not a requirement of holding a licence. + +**A reading is taken on a schedule, not on a request.** The licence-grain poll is a **scheduled +container** ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)) the adapter runs on a cadence; the +session-grain rows are produced as sessions progress, from the transcript the session already writes. +Neither blocks anything: a poll that fails is a logged, retried scheduled run (0053's rule), and a +session whose cost cannot yet be attributed is a row not yet written, never a session refused. + +**Usage is recorded two ways, for two audiences.** Each reading is **emitted as an event** +([ADR 0041](0041-events-are-a-relationship.md)) — an immutable "this was observed at this time" that the +`audit-logger` already records, so the history of what an account did is in the audit trail by default, +under nobody's special arrangement. And the **current** picture — the latest reading per +`(licence, consumer, period, metric)` — is upserted into a **usage context store**, so "how close is +this cap" and "what has this project spent this month" are a query, not a fold over the event log. The +event is the record of what happened; the store is the answer to what is true now. + +**Usage is not a credential, and is recorded in the clear.** The one thing the mesh must not read is the +sealed key ([ADR 0050](0050-model-access-is-vendor-agnostic.md), the refresh-token carve-out aside). A +token count and a cost are not secrets; they are the operator's own operational facts, recorded openly +so they can be queried, audited, and charged against. This is the deliberate opposite of the credential +rule, and stating it prevents a later reader assuming usage inherits the key's secrecy and hiding it +from the person who is paying. + +This is a decision and not a patch because it settles **where a model's usage lives and at what grain**, +which every reader — a bill, a cap alarm, a per-project report — depends on, and because it closes the +half [ADR 0050](0050-model-access-is-vendor-agnostic.md) explicitly left open, reusing the session +([ADR 0026](0026-the-mesh-has-a-session-of-its-own.md)), the schedule +([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)), the event ([ADR 0041](0041-events-are-a-relationship.md)) +and the store ([ADR 0008](0008-a-context-owns-its-store.md)) the mesh already has rather than inventing a +vocabulary beside them. + +### How each claim is checked + +- **A usage row is vendor-neutral and carries the raw beside it.** A unit test constructs an Anthropic + `utilization%` reading and a token/cost reading and asserts both render to + `(licence, consumer, period, metric, value)` with the vendor response preserved in the raw column; + a reader that answers "what did this cost" touches only the normalised columns. +- **The two grains differ only in the consumer.** A unit test records a licence-grain row (consumer = + the module) and a session-grain row (consumer = a session id) for one licence and asserts both are + the same shape and both are returned when the licence's usage is asked for, distinguishable by + consumer. +- **A poll is a scheduled run and its failure is not fatal.** The adapter's usage container declares a + `schedule` ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)); a host test (0053's) already proves a + scheduled step runs on cadence, does not gate, and logs rather than fails on a non-zero run — the poll + inherits this and adds nothing to check. +- **Each reading is an event the audit trail records.** An integration check asserts a usage reading + emits an event that the `audit-logger` receives (it consumes `#`), so the history is present without + the usage module and the audit module knowing about each other beyond the event. +- **The current picture is a query.** A store test upserts two readings for one + `(licence, consumer, period, metric)` and asserts the later replaces the earlier, so "what is true + now" is one row, while the event log keeps both. +- **A static-key vendor with no usage reading records nothing, and that is fine.** A test resolves a + `static-key` model-access consumer whose adapter has no `usage` verb and asserts the licence works and + no usage rows or poll are required — usage is optional, holding a licence is not conditioned on it. +- **Usage is readable in the clear; the key is not.** A test asserts a usage row is stored unsealed and + is returned to an ordinary query, while the licence key remains sealed and absent from the same + surfaces — the deliberate inversion of the credential rule. + +## Consequences + +- **A bill and a cap alarm are both queries.** "What did project X spend this month" reads the + session-grain rows; "how close is account Y to its cap" reads the latest licence-grain metric — both + from the usage store, neither a fold over events or a call to the vendor. +- **The session becomes the unit of cost, which is what it already is.** Because a session is the + consumer, attributing cost needs no new identity — and the worker-granularity gap + ([design 14](../03-DESIGN/01-to-be/14-model-access.md)) surfaces here exactly as it does for access, + to be closed once, for both, when a session id is threaded through. +- **The adapter carries the vendor's usage quirks alone.** Anthropic's `utilization%`, its transcript + shape, a mid-session account switch — all live in the Anthropic adapter; the mesh, the store, and + every reader see only rows. A second vendor adds a second adapter and no new table. +- **Usage history is durable and tamper-evident by reuse, not by a new mechanism.** It rides the event + trail the mesh already keeps, so an operator who wants the whole history has it, and one who wants the + current number has the store — without usage owning either mechanism. +- **The refresh-token carve-out is untouched by this.** Usage is read *from* an authenticated adapter; + it neither holds nor exposes the credential, so the one place the mesh reads what it stores + ([ADR 0050](0050-model-access-is-vendor-agnostic.md)) is not widened by recording what that credential + was spent on. + +## References + +- [ADR 0050](0050-model-access-is-vendor-agnostic.md) — model access is vendor-agnostic; fixes the + usage row shape and the `usage(licence)` adapter verb, and leaves its home open — which this closes +- [ADR 0026](0026-the-mesh-has-a-session-of-its-own.md) — the mesh has a session of its own; a session + is the consumer the finer grain attributes to +- [ADR 0053](0053-a-step-that-runs-on-a-schedule.md) — a scheduled step; the licence-grain poll is one +- [ADR 0041](0041-events-are-a-relationship.md) — events are a relationship; a usage reading is one, and + the audit-logger records it +- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the shape of an event on the wire; the form a + usage reading takes to reach the audit trail +- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store; the current usage picture + lives in one +- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision; usage is a reading of + what that provision was used for +- [03-DESIGN/01-to-be/14-model-access.md](../03-DESIGN/01-to-be/14-model-access.md) — the model-access + design; this fills its usage section and shares its open worker-naming gap +- [03-DESIGN/01-to-be/15-the-agent-session.md](../03-DESIGN/01-to-be/15-the-agent-session.md) — the agent + session; the consumer the session grain is keyed to diff --git a/02-DECISIONS/0055-model-access-is-answered-by-a-licence-or-a-node.md b/02-DECISIONS/0055-model-access-is-answered-by-a-licence-or-a-node.md new file mode 100644 index 0000000..8a994ec --- /dev/null +++ b/02-DECISIONS/0055-model-access-is-answered-by-a-licence-or-a-node.md @@ -0,0 +1,146 @@ +--- +topic: model access +status: accepted +date: 2026-09-07 +deciders: jochen +reconstructed: false +extends: 0050-model-access-is-vendor-agnostic.md +--- + +# 55. Model access is answered by a licence, or by a node that hosts the model + +## Context + +**[ADR 0024](0024-model-access-is-a-provision.md) made model access a provision, and +[ADR 0050](0050-model-access-is-vendor-agnostic.md) fixed what answers it: a record — a licence — that +an adapter turns into a sealed vendor credential.** A consumer requires `model-access`, is put on a +licence, and is delivered a key (a static API key, or an access token a manager refreshes). Every +answer so far has been a credential to reach a vendor's API across the internet. + +**But a model need not come from a vendor. A node in the mesh can host one.** An operator with a GPU +runs Ollama or vLLM, which serves an OpenAI-compatible API on that node. A consumer that wants that +model does not need a vendor credential — it needs the model server's **endpoint**: the base URL and +the model name, and a key only if the server is configured to want one. This is model access answered +by a **node**, not by a record. + +**The mesh already knows how a node answers a provision — it is the ordinary provider/consumer path.** +A provider `provides` a provision at a scope, `serves` its connection facts, and the mesh fills the +consumer's bound facts with the provider's `at`/`port` and each served fact — exactly how a Postgres +consumer learns where its database is. Nothing reserved `model-access` to records: a node offering +`provides: ["model-access"]` resolves through this path, and the resolver already **prefers a local +answer over a licence** — its own comment names the case, "a model the mesh runs itself." So the +capability exists; what is missing is the decision to use it, and the statement of what a node-answer +delivers and where it stops. + +**A node-answer and a record-answer are the same provision with two shapes of answer.** This is the +same move [ADR 0050](0050-model-access-is-vendor-agnostic.md) already made for shapes within the +vendor path (static-key vs refreshable-grant): one provision, more than one way it is answered. A +consumer written against `model-access` should not care whether the model behind it is a vendor's or +the mesh's own — it asks for model access and is given what reaches a model. + +## Considered Options + +1. **A separate provision for the local case (`local-model`, `model-endpoint`).** A node answers that; + `model-access` stays record-only. Rejected: it splits "where my model comes from" into two + provisions a consumer must choose between in its manifest, when the mesh already models a + record-answer and a node-answer to **one** provision. A consumer would have to know, at authoring + time, whether its model will be a vendor's or the mesh's — the exact coupling the provision was + meant to remove. It is the safer implementation (see the limitation below) but the worse interface. + +2. **Model access answered by either a licence or a node, under the one provision.** A consumer + requires `model-access`; the operator answers it with a licence (a vendor) or by assigning a + node that hosts a model. Adopted: one interface, and the answer is an operator's deployment choice, + not a consumer's authoring choice. + +## Decision + +**Model access is one provision answered two ways: by a licence (a record, turned into a sealed vendor +credential by an adapter — [ADR 0050](0050-model-access-is-vendor-agnostic.md)) or by a node that hosts +the model (a provider that serves an endpoint).** A consumer requires `model-access` and is delivered +whichever the operator assigned; it does not name the kind. + +**A node-answer delivers an endpoint, not a credential.** The provider `provides: ["model-access"]` +and `serves` its connection facts — the port it listens on and the model it runs — and the mesh fills +the consumer's bound facts with the provider node's `at`, the served `port`, and the served `model`, +the same way every provider consumer learns where its provider is. The consumer assembles a base URL +(`http://:/v1`) and points an OpenAI-compatible client at it. If the local server wants a +key, the provider mints one the ordinary way (a per-consumer secret, sealed and host-unsealed); if it +does not — the common Ollama case — the consumer lists `model-access` under `binds` and **not** under +`secrets`, and no key is delivered. Secret delivery and fact delivery are already independent, so a +keyless endpoint is expressed by asking for the facts and not a secret. + +**A node-answer uses no adapter.** The adapter registry ([ADR 0050](0050-model-access-is-vendor-agnostic.md)) +is the vendor-credential machinery — accept-and-seal, refresh, usage. A node-hosted model has no vendor +secret to seal; its endpoint is served, and its key (if any) is minted like any provider's. The adapter +is consulted only for the record/vendor answer. So the vendor-agnostic decision is untouched, and the +node-answer adds no vendor logic anywhere. + +**The resolver prefers a local answer.** When a node's own set answers `model-access` — a model the +mesh runs itself — a licence for it is not consulted. This is already the resolver's behaviour and is +made a decision here: a mesh that runs a model uses it, and a licence is the answer for a consumer that +has no local model, not a competitor to one that does. + +**One limitation, stated so it is not found as a bug.** The local-preference above is exact for a +**node-scope** provider co-located with its consumer, and for any node-answer in a mesh that holds no +`model-access` licence. It is *not* yet exact for a **mesh-scope** model server — one node serving the +model to others — **while a licence for `model-access` also exists in the same mesh**: the record pass +that turns a licence into an answer keys on same-node satisfaction and would still demand the licence be +used, double-answering. Until that pass is taught to stand down when a brokered node need already +answers, a mesh-scope local model and a vendor licence must not both answer `model-access` in one mesh. +A node-scope local model has no such constraint. This is named because an unstated limitation is +indistinguishable from a bug, and costs more. + +This is a decision and not a patch because it settles **what may answer model access** — a question +every model-access consumer's meaning depends on — and because it lets the mesh's own hosted models sit +behind the same provision as the vendors', which is what makes "the mesh can run its own model" a +deployment choice rather than a second interface to build against. + +### How each claim is checked + +- **A node answers model access without a licence, and is preferred over one.** A resolver test + assigns a consumer and a module that `provides: ["model-access"]` at node scope on the one node, with + a licence also present, and asserts no resolved need is answered by the record — the local model + answers and the licence is ignored. (This test exists; the decision adopts what it proves.) +- **The consumer is delivered an endpoint, not a credential.** A mesh bed assigns a model-server + provider and a consumer that binds `model-access` and does not list it under `secrets`, and asserts + the consumer's config carries `OPENAI_BASE_URL` built from the provider's served `at`/`port`, and + that no key file was delivered to its secret path. +- **The local endpoint is reachable through what the consumer was given.** The bed makes a request to + the base URL the consumer wrote and asserts the model server answers — the wiring, not a model's + output, is what is proven (the server may be a stub; a real model is not needed to prove the mesh + routed the consumer to it). +- **A node-answer consults no adapter.** A node-answered `model-access` need is resolved with the + vendor registry never read — asserted by the absence of any vendor on a node-answered need and the + ordinary served-facts delivery. +- **The scope limitation holds where stated.** The node-scope case is what the bed and the resolver + test exercise; the mesh-scope-plus-licence collision is recorded here and left for the resolver + change that reconciles the two answer passes, not worked around in a module. + +## Consequences + +- **The mesh can run its own model, and a consumer reaches it through the same `model-access` it uses + for a vendor.** One interface, two answers; a consumer moves between a vendor and a local model by an + operator reassigning its provision, not by a code change. +- **A local model is keyless by default and keyed by the ordinary path when it must be.** Nothing new + is invented for the local server's credential: it either has none, or mints one the way every + provider does. +- **The vendor path is untouched.** Adapters, the refresh carve-out, and usage + ([ADR 0054](0054-model-usage-is-recorded-at-two-grains.md)) are the record answer's business; a + node-answer neither uses nor changes them. A local model that exposes usage would serve it as facts, + not as an adapter's usage verb. +- **The two answer passes meet in one place, and must be reconciled there.** The mesh-scope limitation + is the single point where a node-answer and a record-answer to the one provision can collide; it is + named, and its fix is a resolver change, not a per-module workaround. + +## References + +- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision; this decides a node + may answer it, not only a record +- [ADR 0050](0050-model-access-is-vendor-agnostic.md) — model access is vendor-agnostic; the record + answer and its adapters, which the node answer sits beside and does not use +- [ADR 0054](0054-model-usage-is-recorded-at-two-grains.md) — model usage; a vendor's business on the + record path, served as facts (if at all) on the node path +- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store; a licence is a record in one, + a node-answer needs none +- [03-DESIGN/01-to-be/14-model-access.md](../03-DESIGN/01-to-be/14-model-access.md) — the model-access + design, which this extends with the node answer diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index a599328..5a81355 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -15,7 +15,7 @@ The records run in the order the decisions were taken, oldest first. **Every decision is a record.** There is no ledger and no index file — if a decision is worth recording it is worth a record, and if it is not worth a record it is not recorded -([ADR 0026](0026-every-decision-is-a-record.md)). A "decision" small enough to be one line is +([ADR 0019](0019-how-this-repository-works.md)). A "decision" small enough to be one line is almost always a **rule**, and a rule belongs in [`00-META/how-we-build.md`](../00-META/how-we-build.md), where it is enforced and keeps the incident that earned it. @@ -62,6 +62,91 @@ rather than guessing. ## Index -The index is **generated, not maintained** — run the `hq-status` skill, which reads the -frontmatter of every record. A hand-written index drifts from the folder it describes, and -this one had already done so after a single addition. +**A number identifies a record and never changes.** Records are referenced from outside this +repository — code comments, commit messages — so a number that moves invalidates them silently. +Renumbering once cost 96 references across two code repositories, and that is why the numbers +are now fixed. + +So the folder is in creation order, and **the reading order lives here.** It is generated from +each record's `topic:` and written, because a reader looking at the folder on a forge sees the +folder rather than a command. The objection to a written index is that it drifts — which is +answered by checking it rather than by refusing to write one: + +``` +python3 00-META/checks/index.py --write regenerate +python3 00-META/checks/index.py fail if stale +``` + + + +### What the mesh is + +- **0001** — [The mesh brokers capabilities; nodes host; agents think](0001-mesh-brokers-nodes-host-agents-think.md) +- **0002** — [Nodes communicate over a message broker, not over HTTP](0002-nodes-communicate-over-a-broker.md) +- **0003** — [An agent is a persistent employee, not an instance of a pool](0003-agents-are-persistent-employees.md) + +### Its tiers, from the bottom up + +- **0004** — [A node, and how it joins](0004-a-node-and-how-it-joins.md) +- **0005** — [The node host](0005-the-node-host.md) +- **0006** — [The substrate and the control plane](0006-the-substrate-and-the-control-plane.md) +- **0007** — [Connectivity](0007-connectivity.md) +- **0008** — [A context owns its store, exclusively](0008-a-context-owns-its-store.md) +- **0028** — [The substrate supplies the control plane and nothing else](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) +- **0029** — [A network is a shape, because an action cannot be undone](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) +- **0030** — [Data outlives the mesh that declared it](0030-data-outlives-the-mesh-that-declared-it.md) +- **0031** — [The control plane authenticates nobody, so identity is a module](0031-the-control-plane-authenticates-nobody.md) +- **0033** — [The substrate is a store and a broker](0033-the-substrate-is-a-store-and-a-broker.md) +- **0036** — [Bootstrap ends at a usable mesh, and the first credential comes from a person](0036-bootstrap-ends-at-a-usable-mesh.md) + +### What runs on them, and how it gets there + +- **0009** — [Modules and the graph](0009-modules-and-the-graph.md) +- **0010** — [Delivery](0010-delivery.md) +- **0024** — [Model access is a provision, and a licence is a thing with a name](0024-model-access-is-a-provision.md) +- **0026** — [The mesh has a session of its own, and it is the node session's mechanism](0026-the-mesh-has-a-session-of-its-own.md) +- **0027** — [A provision names what the consumer is coupled to, not the role it plays](0027-a-provision-names-what-the-consumer-is-coupled-to.md) +- **0035** — [One implementation, several surfaces, and what that costs](0035-one-implementation-several-surfaces.md) +- **0038** — [The mesh assigns the port, and a module does not care](0038-the-mesh-assigns-the-port.md) *(proposed)* +- **0040** — [What a module is](0040-what-a-module-is.md) +- **0041** — [Events are a relationship, the lighter sibling of provisioning](0041-events-are-a-relationship.md) +- **0042** — [The shape of an event on the wire](0042-the-shape-of-an-event-on-the-wire.md) +- **0043** — [A module's broker account is scoped by what it emits and consumes](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) +- **0044** — [A public name is provisioned, not registered by hand](0044-a-public-name-is-provisioned-like-any-capability.md) +- **0045** — [A machine's firewall is the sum of what its modules listen on](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) +- **0046** — [A module's configuration is its assignment's, not its manifest's](0046-a-module-configuration-is-its-assignments-not-its-manifest.md) +- **0047** — [A module runs its code as its own process, with its own account](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) +- **0048** — [A provider creates the credential the mesh minted, and seals nothing](0048-a-provider-creates-the-credential-the-mesh-minted.md) +- **0049** — [A consumer's identity is bounded by the tightest backend that must accept it](0049-a-consumers-identity-fits-the-tightest-backend.md) +- **0050** — [Model access is vendor-agnostic, and a vendor is an adapter](0050-model-access-is-vendor-agnostic.md) +- **0051** — [Shared data is the operator's, and a module is granted access to it](0051-shared-data-is-the-operators.md) +- **0052** — [An init step is a container run once to completion, gating what follows](0052-a-step-that-runs-once-before-a-container.md) + +### How it is built + +- **0011** — [Managed files are generated onto nodes and never edited there](0011-managed-files-are-generated-never-edited.md) +- **0012** — [The mesh creates no symlinks — a derived file is a copy](0012-the-mesh-creates-no-symlinks.md) +- **0013** — [Schema and state changes are numbered migrations, in the same language as the code](0013-schema-changes-are-numbered-migrations.md) +- **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md) +- **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md) +- **0016** — [The lab](0016-the-lab.md) +- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)* +- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md) + +### How it is checked + +- **0017** — [A test defends a decision](0017-a-test-defends-a-decision.md) +- **0018** — [A picture of a system is read from the system, never from what asked for it](0018-a-picture-is-read-from-what-runs.md) + +### How we work + +- **0019** — [How this repository works](0019-how-this-repository-works.md) +- **0020** — [The mesh is governed by a constitution, injected where work is decided](0020-the-mesh-is-governed-by-a-constitution.md) +- **0021** — [HQ is the source of the mesh constitution](0021-hq-is-the-source-of-the-constitution.md) +- **0022** — [The constitution absorbs what is already enforced](0022-the-constitution-absorbs-what-is-enforced.md) +- **0023** — [The approval is the checkpoint, not the second pair of hands](0023-approval-is-the-checkpoint.md) +- **0025** — [The design record is read where it is written, never copied to be found](0025-the-design-record-is-read-not-copied.md) +- **0032** — [The local account owns the mesh; a surface delegates to a module](0032-the-local-account-owns-the-mesh.md) *(superseded)* +- **0034** — [The local account owns the mesh, and a web application's login is not that](0034-the-local-account-owns-the-mesh.md) + + diff --git a/03-DESIGN/00-as-is/00-overview.md b/03-DESIGN/00-as-is/00-overview.md index c1d6a0d..a09fe84 100644 --- a/03-DESIGN/00-as-is/00-overview.md +++ b/03-DESIGN/00-as-is/00-overview.md @@ -4,9 +4,9 @@ status: implemented code: [hal] updated: 2026-08-23 decisions: - - 02-DECISIONS/0001-nodes-communicate-over-a-broker.md - - 02-DECISIONS/0002-everything-is-a-module.md - - 02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md + - 02-DECISIONS/0002-nodes-communicate-over-a-broker.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0006-the-substrate-and-the-control-plane.md --- # The mesh as it stands @@ -28,11 +28,11 @@ onto it and can be regenerated. containerised service is a module. A set of capabilities with no service behind them is a module. A bare marker whose whole content is that a node has it is a module. The mesh's own components are modules on exactly the same terms as everything else it carries -([ADR 0002](../../02-DECISIONS/0002-everything-is-a-module.md)). +([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)). **An agent** is a participant. Some agents are human. What differs is modality — how the agent acts — and not category: both hold identity, both act, both accumulate memory -([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)). +([ADR 0003](../../02-DECISIONS/0003-agents-are-persistent-employees.md)). ## Where truth lives @@ -40,10 +40,10 @@ The repository defines **what exists**: the modules, what each declares, how eac The mesh database defines **what runs where**: which node is assigned which module, at which selection, with which overrides, plus the settings every node reads. No node-to-module mapping -is ever committed ([ADR 0003](../../02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md)). +is ever committed ([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). Everything on a node's disk is **derived** from those two, and is regenerated rather than -edited ([ADR 0004](../../02-DECISIONS/0004-managed-files-are-generated-never-edited.md)). A node that +edited ([ADR 0011](../../02-DECISIONS/0011-managed-files-are-generated-never-edited.md)). A node that loses its database keeps running from a local cache, which is deliberate and has the obvious cost: the cache carries no indication of its own age. @@ -51,7 +51,7 @@ cost: the cache carries no indication of its own age. Nothing dials a node. Every node dials the broker outbound, owns an exchange named for itself, and consumes from its own request queue -([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)). Three message shapes carry +([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)). Three message shapes carry everything: requests expecting a reply, commands instructing that a stage of work be done, and events stating that something happened. @@ -65,9 +65,9 @@ goes to where the capability is. A push to the forge is the only trigger. What follows is three silos with deliberately different cardinality: compile once, package and upload once, then install-configure-start- verify **on every assigned node** -([ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)). What travels between +([ADR 0010](../../02-DECISIONS/0010-delivery.md)). What travels between build and node is a self-contained build output, so a deploy is extract-and-run and touches no -network ([ADR 0013](../../02-DECISIONS/0013-an-artifact-is-build-output.md)). +network ([ADR 0010](../../02-DECISIONS/0010-delivery.md)). Modules are resolved into dependency levels and a level completes before the next begins, so a module always builds against its dependencies as they were just published. @@ -78,7 +78,7 @@ A module declares what it **provides** and what it **requires**. The mesh satisf requirement: it creates the resource, generates the credential, records the grant, and writes the values where the module will read them. The module never learns which node its database lives on, and nobody ever writes a credential by hand -([ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md)). +([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)). This is the property the mesh's whole shape rests on, and it is why provisioning is treated as a core concern rather than as plumbing. @@ -92,7 +92,7 @@ named for a feature the module does not declare, a stage that reported it had di message rather than that the effect happened, a package that 404ed from every mirror while the job went green. -[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) is the response, and it is applied +[ADR 0010](../../02-DECISIONS/0010-delivery.md) is the response, and it is applied instance by instance rather than enforced by a mechanism. New instances are still being found. That is an as-is fact, not a criticism: it is the single most useful thing to know about this system before changing it. diff --git a/03-DESIGN/00-as-is/01-mesh-and-transport.md b/03-DESIGN/00-as-is/01-mesh-and-transport.md index 7416575..6e4848e 100644 --- a/03-DESIGN/00-as-is/01-mesh-and-transport.md +++ b/03-DESIGN/00-as-is/01-mesh-and-transport.md @@ -4,8 +4,8 @@ status: implemented code: [hal] updated: 2026-08-23 decisions: - - 02-DECISIONS/0001-nodes-communicate-over-a-broker.md - - 02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md + - 02-DECISIONS/0002-nodes-communicate-over-a-broker.md + - 02-DECISIONS/0006-the-substrate-and-the-control-plane.md --- # The mesh and its transport diff --git a/03-DESIGN/00-as-is/02-modules-and-manifests.md b/03-DESIGN/00-as-is/02-modules-and-manifests.md index 4cc8875..4c50770 100644 --- a/03-DESIGN/00-as-is/02-modules-and-manifests.md +++ b/03-DESIGN/00-as-is/02-modules-and-manifests.md @@ -4,9 +4,9 @@ status: implemented code: [hal] updated: 2026-08-23 decisions: - - 02-DECISIONS/0002-everything-is-a-module.md - - 02-DECISIONS/0006-schema-changes-are-numbered-migrations.md - - 02-DECISIONS/0007-no-npm-workspace.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0013-schema-changes-are-numbered-migrations.md + - 02-DECISIONS/0014-no-npm-workspace.md --- # Modules, manifests and features @@ -81,7 +81,7 @@ recorded in the knowledge base; both presented as "the change did not apply" wit ## Dependencies between modules Modules depend on each other, above all on the shared library they all build against. There is -**no workspace** ([ADR 0007](../../02-DECISIONS/0007-no-npm-workspace.md)): each module is a standalone +**no workspace** ([ADR 0014](../../02-DECISIONS/0014-no-npm-workspace.md)): each module is a standalone package consuming published dependencies, including the mesh's own. The pipeline resolves modules into dependency **levels** and completes a level before starting @@ -96,7 +96,7 @@ since (see A module that owns state owns its migrations: numbered, written in the module's own language, compiled with it, frozen once they have run anywhere, and idempotent so that re-running is safe -([ADR 0006](../../02-DECISIONS/0006-schema-changes-are-numbered-migrations.md)). +([ADR 0013](../../02-DECISIONS/0013-schema-changes-are-numbered-migrations.md)). Two kinds exist and the distinction matters: migrations against the module's **own** local state, and migrations against a **provisioned** resource, which run on the node that consumes diff --git a/03-DESIGN/00-as-is/03-provisioning.md b/03-DESIGN/00-as-is/03-provisioning.md index aa199c8..f3e6524 100644 --- a/03-DESIGN/00-as-is/03-provisioning.md +++ b/03-DESIGN/00-as-is/03-provisioning.md @@ -4,8 +4,8 @@ status: implemented code: [hal] updated: 2026-08-23 decisions: - - 02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md - - 02-DECISIONS/0004-managed-files-are-generated-never-edited.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0011-managed-files-are-generated-never-edited.md --- # Provisioning diff --git a/03-DESIGN/00-as-is/04-delivery.md b/03-DESIGN/00-as-is/04-delivery.md index 0f55da7..9c8b891 100644 --- a/03-DESIGN/00-as-is/04-delivery.md +++ b/03-DESIGN/00-as-is/04-delivery.md @@ -4,9 +4,9 @@ status: implemented code: [hal] updated: 2026-08-23 decisions: - - 02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md - - 02-DECISIONS/0013-an-artifact-is-build-output.md - - 02-DECISIONS/0008-a-failed-step-fails-the-job.md + - 02-DECISIONS/0010-delivery.md + - 02-DECISIONS/0010-delivery.md + - 02-DECISIONS/0010-delivery.md --- # Delivery — from a push to a running node @@ -31,7 +31,7 @@ merge that created no pipeline, and nothing said so**. ## Three silos Cardinality is the whole point, and the three differ -([ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)): +([ADR 0010](../../02-DECISIONS/0010-delivery.md)): | Silo | Runs | Where | Does | |---|---|---|---| @@ -50,7 +50,7 @@ later stage runs. ## The artifact The artifact is **build output** — compiled and bundled with its dependency graph inlined — -never a filtered copy of source ([ADR 0013](../../02-DECISIONS/0013-an-artifact-is-build-output.md)). +never a filtered copy of source ([ADR 0010](../../02-DECISIONS/0010-delivery.md)). A deploy is extract-and-run and touches no network. The consequence is the whole cost of the decision: **anything not in the build output does not diff --git a/03-DESIGN/00-as-is/05-runtime-and-installation.md b/03-DESIGN/00-as-is/05-runtime-and-installation.md index 3db208d..cf06182 100644 --- a/03-DESIGN/00-as-is/05-runtime-and-installation.md +++ b/03-DESIGN/00-as-is/05-runtime-and-installation.md @@ -4,8 +4,8 @@ status: implemented code: [hal] updated: 2026-08-23 decisions: - - 02-DECISIONS/0002-everything-is-a-module.md - - 02-DECISIONS/0011-the-installer-owns-linking.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0012-the-mesh-creates-no-symlinks.md --- # The node runtime, and how a node comes into being @@ -38,7 +38,7 @@ suggestive word in the system names the node runtime, and the component whose ma "mesh messaging" is documented elsewhere as the interactive runtime. Anatomy makes attractive names and poor boundaries. -[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) replaces this with names +[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) replaces this with names taken from what each part owns. Until then, this is the vocabulary in the code. ## Starting a module @@ -57,14 +57,14 @@ outstanding local migrations, create data directories with the right ownership, service under supervision. **The installer is the only thing that creates a link** ([ADR -0011](../../02-DECISIONS/0011-the-installer-owns-linking.md)). It reconciles rather than assumes: a +0011](../../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md)). It reconciles rather than assumes: a missing link is created, a stale one repointed, and a real file found where a link belongs is adopted into the node's override area and replaced. Nothing else — not a hook, not a fix, not a person debugging — creates one. That is the as-is. The intent is to remove linking altogether and derive a real file instead, which the reconciliation machinery already makes possible -([ADR 0018](../../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md), proposed). What is described above +([ADR 0012](../../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md), proposed). What is described above is what runs today. ## Supervision diff --git a/03-DESIGN/00-as-is/06-configuration-and-secrets.md b/03-DESIGN/00-as-is/06-configuration-and-secrets.md index 8372519..5de0385 100644 --- a/03-DESIGN/00-as-is/06-configuration-and-secrets.md +++ b/03-DESIGN/00-as-is/06-configuration-and-secrets.md @@ -4,8 +4,8 @@ status: implemented code: [hal] updated: 2026-08-23 decisions: - - 02-DECISIONS/0004-managed-files-are-generated-never-edited.md - - 02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md + - 02-DECISIONS/0011-managed-files-are-generated-never-edited.md + - 02-DECISIONS/0009-modules-and-the-graph.md --- # Configuration and secrets @@ -17,7 +17,7 @@ files is **generated**. A managed file is derived from the mesh database. A synchroniser rewrites it when the values behind it change. The write path is the mesh operation that owns the value; the file is an -output ([ADR 0004](../../02-DECISIONS/0004-managed-files-are-generated-never-edited.md)). +output ([ADR 0011](../../02-DECISIONS/0011-managed-files-are-generated-never-edited.md)). An edit to a managed file survives until the next synchronisation and is then overwritten silently, taking whatever it was fixing with it — bringing back the bug the edit had removed, @@ -61,7 +61,7 @@ are both left behind. Configuration is additive in practice, whatever the manife Generated secrets are produced by the mesh, never authored. Provisioned credentials arrive as database overrides written by the provisioner and are marked as such, so they can be distinguished from a deliberate override and cleaned up when the grant is removed -([ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md)). +([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)). Nothing in the repository contains a credential. The repository has no per-node content at all, which is what makes that guarantee structural rather than a matter of care. diff --git a/03-DESIGN/00-as-is/07-knowledge.md b/03-DESIGN/00-as-is/07-knowledge.md index b7d14ed..fb09147 100644 --- a/03-DESIGN/00-as-is/07-knowledge.md +++ b/03-DESIGN/00-as-is/07-knowledge.md @@ -40,7 +40,7 @@ owning approval and promotion at the boundary. Proposals to edit are reviewed ra applied. This is where the mesh's **governed** documents live, including the constitution injected into -design sessions ([ADR 0009](../../02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md)). +design sessions ([ADR 0020](../../02-DECISIONS/0020-the-mesh-is-governed-by-a-constitution.md)). ## Why both diff --git a/03-DESIGN/00-as-is/08-agents-and-work.md b/03-DESIGN/00-as-is/08-agents-and-work.md index d6d3877..c51f5ab 100644 --- a/03-DESIGN/00-as-is/08-agents-and-work.md +++ b/03-DESIGN/00-as-is/08-agents-and-work.md @@ -4,8 +4,8 @@ status: implemented code: [hal] updated: 2026-08-23 decisions: - - 02-DECISIONS/0012-agents-are-persistent-employees.md - - 02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md + - 02-DECISIONS/0003-agents-are-persistent-employees.md + - 02-DECISIONS/0020-the-mesh-is-governed-by-a-constitution.md --- # Agents and work @@ -17,7 +17,7 @@ model they run under is the employee model, not a worker pool. An agent is a singular named identity with a home node, a workspace on that node, accumulating memory, and an explicit lifecycle -([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)). +([ADR 0003](../../02-DECISIONS/0003-agents-are-persistent-employees.md)). | Property | Meaning | |---|---| @@ -45,7 +45,7 @@ Both hold identity, both act, both accumulate memory. The mesh does not currently record modality completely. Which user, on which node, a human agent acts as is **required by the model and not stored** — an open question carried over from -[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md). +[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md). ## Work @@ -70,7 +70,7 @@ template that names the phases. This is where governance meets execution. The constitution is injected into every eligible meeting turn — agents do not fetch it, it arrives — and a check phase verifies the meeting's output against it before the meeting may proceed -([ADR 0009](../../02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md)). A named violation +([ADR 0020](../../02-DECISIONS/0020-the-mesh-is-governed-by-a-constitution.md)). A named violation blocks progress. Meeting turns run on the orchestrator's node regardless of where the participating agents are @@ -84,5 +84,5 @@ integrate through the record, never through a shared schema* — being violated own largest component, and it is the reason work that belongs to one domain keeps having to be implemented in another. -[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) dissolves that arrangement. +[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) dissolves that arrangement. Until it does, this is the shape. diff --git a/03-DESIGN/00-as-is/09-interfaces-and-observability.md b/03-DESIGN/00-as-is/09-interfaces-and-observability.md index f47665d..6f08b04 100644 --- a/03-DESIGN/00-as-is/09-interfaces-and-observability.md +++ b/03-DESIGN/00-as-is/09-interfaces-and-observability.md @@ -4,8 +4,8 @@ status: implemented code: [hal] updated: 2026-08-23 decisions: - - 02-DECISIONS/0001-nodes-communicate-over-a-broker.md - - 02-DECISIONS/0008-a-failed-step-fails-the-job.md + - 02-DECISIONS/0002-nodes-communicate-over-a-broker.md + - 02-DECISIONS/0010-delivery.md --- # Interfaces and observability diff --git a/03-DESIGN/00-as-is/10-module-catalogue.md b/03-DESIGN/00-as-is/10-module-catalogue.md index 899aa78..57de760 100644 --- a/03-DESIGN/00-as-is/10-module-catalogue.md +++ b/03-DESIGN/00-as-is/10-module-catalogue.md @@ -4,16 +4,16 @@ status: implemented code: [hal] updated: 2026-08-23 decisions: - - 02-DECISIONS/0002-everything-is-a-module.md - - 02-DECISIONS/0010-applications-live-in-their-own-repository.md - - 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0015-applications-live-in-their-own-repository.md + - 02-DECISIONS/0009-modules-and-the-graph.md --- # The catalogue, and what its shape says The catalogue holds **124 modules**. Thirty-three belong to the mesh's own domain; the other ninety-one run *on* the mesh rather than being *of* it -([ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)). +([ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md)). The count is not the finding. The **shape** is. @@ -55,11 +55,11 @@ connectivity is made four times. the unit of one piece of software, because that is the only granularity the module system offers. -This is the same failure [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) +This is the same failure [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) names for the platform core — *boundaries drawn by deployment accident rather than by domain* — appearing outside it, at four times the scale. The core is being recomposed; the flat level is addressed in principle by -[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md), which +[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md), which deliberately does not yet settle the domain list. ## Where the shape came from @@ -90,7 +90,7 @@ never stated as assumptions — they were just how the thing already worked. **This is the most useful single fact for anyone changing the catalogue**, and it is why the linking principle in particular reads as a deliberate architectural choice when it is an -inheritance. See [ADR 0018](../../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md), whose case +inheritance. See [ADR 0012](../../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md), whose case this strengthens: the argument for links was never made *for a mesh*. It also explains the measurement in @@ -108,7 +108,7 @@ which is what makes dogfooding structural rather than a discipline, and what mak module out of the repository safe. **Placement is already decided.** A standalone application belongs in its own repository -([ADR 0010](../../02-DECISIONS/0010-applications-live-in-their-own-repository.md)), and reviewers reject +([ADR 0015](../../02-DECISIONS/0015-applications-live-in-their-own-repository.md)), and reviewers reject it in the monorepo. The catalogue's flat level is not a dumping ground by policy; it is one by history. diff --git a/03-DESIGN/00-as-is/11-the-lab.md b/03-DESIGN/00-as-is/11-the-lab.md index 30b977d..9ebff58 100644 --- a/03-DESIGN/00-as-is/11-the-lab.md +++ b/03-DESIGN/00-as-is/11-the-lab.md @@ -4,11 +4,11 @@ status: implemented code: [mesh-lab] updated: 2026-08-25 decisions: - - 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md - - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md - - 02-DECISIONS/0031-the-lab-provides-the-underlay.md - - 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md - - 02-DECISIONS/0033-a-router-is-scenery-not-a-node.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md --- # The lab, as it stands @@ -63,7 +63,7 @@ is built. **The drawing was never designed.** `diagram` renders a scenario as draw.io, from the declaration or from the running instance, and it exists because it was asked for during the -build. It has tests and a decision record ([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md), +build. It has tests and a decision record ([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md), proposed) but no document in the to-be layer. It is recorded here because it runs, not because it was planned. @@ -117,7 +117,7 @@ and snapshots roughly 76× slower, which does not make the lab slow, it makes it `npm run check` — typecheck over source *and* tests, then the offline suite, then integration against a real hypervisor. Mocking the hypervisor is forbidden -([ADR 0034](../../02-DECISIONS/0034-a-test-defends-a-decision.md), proposed): a test that fakes +([ADR 0017](../../02-DECISIONS/0017-a-test-defends-a-decision.md), proposed): a test that fakes the system under integration asserts that the fake behaves as expected. Integration tests **skip with a reason** on a machine that cannot raise scenarios, rather than diff --git a/03-DESIGN/01-to-be/00-work-breakdown.md b/03-DESIGN/01-to-be/00-work-breakdown.md index e39762e..6db2583 100644 --- a/03-DESIGN/01-to-be/00-work-breakdown.md +++ b/03-DESIGN/01-to-be/00-work-breakdown.md @@ -1,158 +1,430 @@ --- layer: to-be status: designed -code: [hal] -updated: 2026-08-23 -decisions: [02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md] +code: [] +updated: 2026-09-01 +decisions: + - 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md --- -# Work breakdown — the decomposition +# Work breakdown — replacing what provisions the mesh -How [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) gets built, in what order, and where a human must look. +*Rewritten 2026-08-31. The previous version planned a decomposition of the existing system in +place: extract contexts, convert modules to declared features, shrink its shared library. That is +not what is being done — a replacement is being built beside it, and the old plan's Phase 0 was +the only part that survived contact with it. So the document that was supposed to say what happens +next had been describing work on a system being retired.* -Ordering is not preference. Each phase removes a constraint the next one needs gone. +## The goal, in one sentence ---- +**Modules move to the new mesh one at a time, until the old registry can be switched off.** + +Everything below is ordered by what that requires. Nothing here is a rewrite of the old system; +its modules are the input. + +## Phase 0 — a mesh that runs — **done** + +Not *the code exists*. Twenty-two assertions on real machines in the lab, each confirmed to fail +when the behaviour is removed ([ADR 0016](../../02-DECISIONS/0016-the-lab.md), +[ADR 0017](../../02-DECISIONS/0017-a-test-defends-a-decision.md)). + +| what is proven | | +|---|---| +| **a mesh comes into being** | a bare machine becomes one; others join with nothing but a token | +| **credentials** | delivered to both ends with the mesh holding neither; rotated so the old one stops working | +| **declarations survive reality** | a stopped machine is waited for; one that fell behind catches up unnamed; unassigning takes away exactly what it should; what the mesh says nothing about is left alone | +| **failure is legible** | a machine that cannot do what it was told is named, with why | +| **the mesh runs itself** | its own artifact store, and a builder that is a module the mesh assigns | +| **names and reachability** | internal names, wildcards under a machine, containers reaching other machines, certificates the mesh issued, filtering that matches exactly what was declared | +| **delivery** | a new commit reaches a machine already running the old one | +| **model access** | answered by a record, with a key the mesh cannot read | + +**What Phase 0 does not prove, and it is the important sentence in this document:** every module +exercised above was written to test the mechanism. **No module from the existing system has ever +run on this.** The vocabulary was shaped by the things used to test it — the same fault as a +fixture agreeing with the code it checks +([`04-ISSUES/005`](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)), at the +scale of a design. + +## Phase 1 — the vocabulary a real module needs + +Found by taking real modules and asking what they would require. Each is a gap in what can be +*expressed*, not a defect in what is built. + +| # | task | done when | +|---|---|---| +| ~~1.1~~ | ~~An **object-store provision**~~ — **done 2026-08-31**, and it needed no change to the mesh: see below | seven assertions against a real store | +| ~~1.2~~ | ~~**A session as a consumer of a licence**~~ — **done 2026-08-31**, and it also needed no change: see below | two sessions on one machine, different licences, each its own key | +| ~~1.3~~ | ~~A **network** shape, and ordering within a module~~ — **done 2026-08-31** | the shape is created and removed; ordering was already there, and is now asserted | +| 1.4 | **Public certificate issuance** — **built; one gap** | ordering, the challenge and issuance are proven against a real authority; **collecting the issued certificate is not** ([`04-ISSUES/020`](../../04-ISSUES/020-a-certificate-is-issued-and-never-collected/00-report.md)) | + +**1.3 and 1.4 block later ones** and are listed now so they are not met as surprises. 1.3 is what +a mail system needs and nothing else so far does. + +**Phase 1 is closed with 1.4 partly open**, deliberately. Two of its four items needed no code at +all; the network shape was built; and certificates are configured correctly, order correctly, and +are issued correctly — the client does not collect what the authority issued, against a server +that exists to be a test server. That is filed rather than chased, because the remainder may say +nothing about a real authority and the next thing to learn comes from moving a module rather than +from a fourth lab run. + +**Checkpoint:** each is demonstrated in the lab before the module needing it is attempted. + +### 1.2, and the same surprise twice + +**A binding is per module per machine, and the two sessions are two modules** — the same mechanism +in different context roots, and a context root is what a module delivers. So `(node, module)` +already names them apart, and nothing needed adding. +[`14-model-access.md`](14-model-access.md) had called per-module-per-machine *a step toward it and +not it*, which is true of a **worker** — many run on one machine from one module — and not true of +a session, of which there is one per node and one for the mesh. + +### 1.3, and the first one that needed building + +**Ordering was already there** — the apply loop sorts nothing, so a module says *this before that* +by writing it first. Untested until now, and the kind of property a later change breaks silently. +Worth separating from readiness: a container started is not a container ready, and nothing waits. +What needs something *usable* retries, which is what both provisioners do and is the better answer +anyway, because a dependency can restart long after everything was applied. + +**The network was a real gap, and the first thing in Phase 1 that needed a decision.** Adding a +shape widens what a compromised control plane can express, so +[ADR 0029](../../02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) +records why this one is worth it: an `action` could create a network and **nothing could ever +remove it**, because an action leaves no footprint the host can undo. The vocabulary is nine. + +**Three tasks in a row that were already possible.** Both were written from the design rather than +from the code, which is the review's finding arriving in the plan: *a claim here is counted, not +reasoned.* The remaining Phase 1 items should be checked against the code before being started, +not after. + +### 1.1, and what it turned out to be + +*Done 2026-08-31. Worth recording because the task was not the one written down.* + +**The control plane special-cases nothing.** `provides`, `requires`, `contributes` and `grants` +are name-agnostic — asking for a bucket needed no change to the mesh at all. What was missing was +a provider, and the last step where something on the machine turns a delivered secret into a key +that works. So "add an object-store provision" was never mesh work. + +The provision is `s3-bucket`: a consumer's code is written against the S3 API and swapping one +store for another does not break it, so by +[ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md) the name +says the protocol. A database is the other case, and names the engine. + +**One assertion here that a database does not need.** One PostgreSQL server holds separate +databases and the product enforces the boundary; one object store holds every bucket behind one +endpoint, so *a consumer cannot reach another consumer's bucket* is a policy somebody wrote — and +a policy granting everything would pass every other test. **What is asserted is what the policy +does not say.** + +## Data is the constraint, and it outranks the order below + +*2026-08-31.* The modules being converted run live services — identity, mail — and **the data must +survive every step**. A data folder may move; it may never be lost. + +**One thing was found by asking this and is fixed** +([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)): the host deleted +a directory and everything under it when the directory stopped being declared, which happens when +a module is unassigned or a manifest is edited to move a data folder — the exact operation this +plan needs. A directory holding anything the mesh did not put there is now kept and reported. + +**That is not a backup and must not be read as one.** It stops the mesh destroying data. It does +nothing about a disk, a mistaken command, or a service corrupting its own store. + +**So the rule for every step below:** the data is copied, the copy is verified by reading it back +through the service that owns it, and only then does anything point at the new location. Never +moved and then checked. **A backup nobody has restored is a belief, not a copy.** + +## A module is adopted with the credentials it already has + +*2026-08-31.* **Nothing is rotated during the conversion.** A service being adopted keeps the +password it is already using, because minting a new one is how a running service stops being able +to reach its own database in the middle of a migration. + +The mesh has both paths and this needs the second: + +| | | +|---|---| +| **generate** | a new secret, sealed to both ends. What a *new* module gets | +| **accept** | a value supplied from outside, sealed, plaintext discarded. **What an adopted module gets** | + +**Rotation is a separate act, afterwards, once everything works.** The machinery for it is built +and proven — a credential moving at both ends with the old one ceasing to work — and it is exactly +the sort of thing to do deliberately on a quiet afternoon rather than as a side effect of moving a +service between systems. + +**So there is a step before any of this: read the current environment out of the old system**, because +adoption means supplying those values and they live in its files today. + +**And there is a failure worse than losing data, which is likelier.** A database image consumes its +password environment variable **only when its data directory is empty**. Everything here keeps its +data on a persistent directory, so the role holds whatever password it was created with, for ever. +Regenerate that variable and the application moves on while the database does not — permanently, +because nothing reconciles it. Eight modules are in that state today, working only because nobody +has regenerated their credential since their data directory was created. + +*Where the detail lives:* this is operational and names machines, so it is in the mesh's own +knowledge base rather than here — `migration/where-service-data-lives`, which surveys where every +service's data actually sits and what each stop or removal would cost, and +`troubleshooting/db-password-frozen-at-first-init` for the lockout itself. **This document says the +rule; those say the specifics.** + +*Corrected 2026-08-31 — an earlier version of this paragraph made that sound more dangerous than it +is.* A sealed secret is not unreadable; it is sealed **to the node**, which holds the private half +and writes the plaintext into the module's own file. The value is there, on the machine, as an +ordinary file. What does not exist is a way to ask *the mesh* what a secret is, and there is no +reveal command, because a mesh that can reveal a secret is a mesh that holds one. + +## Where it starts, and what that costs + +**On the node holding all the production data**, because that is where the services being +converted actually are. + +Recorded plainly rather than argued with: this is the highest-risk order available. Everything +proven so far was proven on machines that could be destroyed and raised again, and the first real +exercise of the conversion will be on the one machine where a mistake is not recoverable. Nothing +about the lab work transfers automatically — a scenario proves the mechanism, not the state on +that machine. + +**What makes it survivable is preparation rather than caution**: a restored backup before the +first step, one service at a time, and the previous arrangement left standing until the new one +has been read back. None of that is slower than the alternative, because the alternative includes +losing something. + +## How the two systems hand over + +*2026-08-31.* **The old system's brain is switched off; its services keep running.** + +Not a migration and not a period of dual control. The old control plane — provisioning, the +coordinator, the pipeline, the things that *decide* and *write* — is stopped. Every workload it +was managing goes on running exactly as it is, because nothing is managing it. Then the new mesh +takes ownership of them one at a time. + +**Nothing is ever unassigned in the old system.** Unassigning is how it removes things, and +removing is how data is lost. The old system is never asked to take anything away; it is asked to +stop having opinions. + +| | | +|---|---| +| **stopped, and disabled** | provisioning, the coordinator, environment and configuration sync, the pipeline — anything that decides or writes a file | +| **left alone entirely** | the units running the actual services: identity, mail, databases, the forge. They keep serving throughout | +| **never used** | unassign, remove, delete — any operation whose job is to take something away | + +**Disabled, not merely stopped**, and this is the part that is easy to get wrong: those units are +enabled, so stopping them lasts until the machine reboots. A reboot mid-conversion would bring the +old control plane back and it would resume regenerating managed files underneath the new one — +which is the one situation where two systems really would be fighting over the same machine. + +**A service left running with nothing managing it is the safe state.** It has its data, its +configuration is already on disk, and nothing is going to change either. That is the whole trick: +the risk in a conversion is in the *managing*, not in the *running*. + +**A brief interruption is acceptable. Losing data is not.** Where those two trade against each +other, the interruption wins every time — a service can be restarted, and there is no operation +that un-deletes a mail spool. + +**The new host cannot remove what it did not put there.** Orphans are per-origin, so it only ever +removes resources it recorded itself. Services it has never been told about are not orphans to +it — they are simply not its business, which is what makes taking ownership one module at a time +safe. + +## Phase 2 — the first real module + +| # | task | done when | +|---|---|---| +| 2.1 | Port an **object store** module | it runs on the new mesh, serves a bucket to another module, and its credential rotates | +| 2.2 | Copy the data, and read it back through the service that owns it | the new location answers with what the old one holds | +| 2.3 | Point one dependent at it, old arrangement left standing | something real reads and writes through the new mesh's copy | + +**Checkpoint, and it is a human one:** it runs for a week before anything else moves. The point of +going first is to find what Phase 1 missed, and a week is roughly how long that takes to show. + +## Phase 3 — the modules that prove the shape + +Each exercises something the first one does not. + +| # | task | proves | +|---|---|---| +| 3.1 | An **identity provider** | a module that is itself a provider — the provides/requires chain, with consumers requiring it | +| 3.2 | A **forge** | a port claim against the machine's own daemon, and a module wanting both a database and an object store | +| 3.3 | A **mail system** | several containers as one module, a private network between them, and names that are not one-per-node | + +**3.3 is the hardest thing in this document** and is deliberately last. If the declaration +language turns out to be insufficient, it says so here. + +### Where Phase 3 actually stands — *2026-09-01* + +All three have manifests. All three parse, resolve and plan. **None of them can start**, and the +two reasons are both filed rather than guessed at. + +The **vocabulary held**. Nothing in 3.1–3.3 turned out to need a new shape: the identity provider, +the forge and the mail system are all expressible with what exists, including the mail system's +several containers on a private network — which was the one expected to break it. That is the +question this phase was designed to answer, and the answer is yes. + +What did not hold was underneath the vocabulary: + +- **[`022`](../../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md)** + (fixed) — a credential belonged to a machine, so a node running several modules against one + database could not be planned. The refusal was loud on the provider and silent on the consumer. +- **[`023`](../../04-ISSUES/023-a-consumer-cannot-build-a-connection-string/00-report.md)** (open) + — a consumer gets its password and still cannot connect: the user name is invented by the + provisioner and recorded nowhere, and the bound values cannot reach a configuration file. + +A third fault was in the manifests themselves rather than the design: each declared a secret at a +path named `.env` and read it as one, when a sealed file holds a password and nothing else. They +parsed and resolved and could never have worked, which is what a manifest checked only by a parser +buys. Two tests now refuse both halves of it. + +### 3.1 needs a program, not a decision — *2026-09-01* + +023 is fixed, and with it the two design faults are gone. What stands between 3.1 and a running +identity provider is now one concrete thing: **the realm provisioner does not exist.** + +Its manifest named an image — `mesh-provision-keycloak` — that nothing builds and no program +backs. That has been removed rather than left standing, because a manifest describing a program +nobody wrote is the same mistake as the credential files that could never be read: it parses, it +resolves, and it could never work. + +So Keycloak's manifest now says what is true today — a server the mesh runs, with its database +and its admin credential, both reaching it in a shape it can read. It no longer claims to provide +`oidc-client`, which means a consumer asking for one is **refused by name at plan time** rather +than resolving cleanly and waiting for a client nothing will create. + +The provisioner is the same shape as the two that exist: it reads what the mesh granted and +reconciles a realm and a client per consumer. **It should be written against a real Keycloak in +the lab**, not from the API documentation — the object store's took three corrections that only a +running server produced. + +The forge (3.2) and the mail system (3.3) need no provisioner and are not blocked on this. + +### And they could not have run anyway — *2026-09-01* + +Every one of the five named a container image that does not exist: sixty-four zeros where a digest +belongs, eighteen times over +([`025`](../../04-ISSUES/025-a-module-must-pin-a-digest-and-nothing-produces-one/00-report.md)). +They parsed, resolved and composed into a declaration a host accepts, and every one would have +stopped on the machine at the moment of fetching. + +Nothing caught it because nothing could. A host checks the *shape* of a reference and no more — +verifying a digest exists means reaching a registry, which is the one thing a host must never have +to do. The refusal now sits where a declaration is composed instead, which is the last moment +before a machine sees one. + +Twelve are pinned to real images. Two further faults surfaced only by pinning for real: the mail +system's seven images named repositories that **do not exist**, because it publishes to a +different registry than assumed, and one of the seven had been renamed upstream. + +**The forge now runs**, on a database another module provides, with a password it did not choose +and a connection string it could not have written. That is the first of these descriptions to be +started rather than planned, and it exercises everything the credential work added. + +What is still missing is the mechanism: nothing turns a tag into a digest as part of the mesh's +own work, so it was done by hand. Asking a registry takes about a second and pulls nothing, which +removes the main argument for leaving it undone. + +## The conversion is done by hand, and that is a decision + +*2026-08-31.* **Moving from the current system to this one is a person at a command line, working +through it.** Not a migration program, not a converter, not a period of dual-writing. + +**What that removes from this plan is larger than what it adds.** Nothing below needs an importer, +a translation layer, a compatibility shim, or a mechanism for keeping two systems agreeing while +both are live — and every one of those is a thing somebody would otherwise reasonably build, use +once, and maintain for a year. The modules are the input; a person reads what one does today and +writes what it declares tomorrow. + +**It also changes what "safe" means for the system being retired.** A fix to it has to be safe on +its own, because there is no careful rollout to sequence it into: the thing is being switched off +by hand, not managed into retirement. A change needing three steps in the right order is a change +that will be half-applied. + +**And it is why the checkpoints below are weeks rather than gates.** Nothing enforces the order — +a person does — so the value of the sequence is entirely in what each step teaches before the next +one starts. + +## Phase 4 — switch the old registry off + +| # | task | done when | +|---|---|---| +| 4.1 | Move the remainder, by hand, a module at a time | nothing is assigned in the old system that is not assigned in the new one | +| 4.2 | The old one authoritative for nothing | a change to any module goes through the new mesh only | +| 4.3 | Switch it off | it is stopped, and nothing notices | + +**4.3 is a day's work and the phases above it are not.** Naming it as a phase is what stops it +being mistaken for the goal. + +## Sequencing + +- **1 before 2.** Attempting a module without the vocabulary it needs produces a workaround, and a + workaround in a manifest is a design decision taken by whoever was in a hurry. +- **2 before 3, with the week.** Moving three modules before running one is how three modules + acquire the same defect. +- **3.3 last.** It is the only one that may send work back into the declaration language. +- **4 cannot start early, and there is no partial credit.** A registry still authoritative for one + module is still running. + +## How this list is kept true + +*This section exists because the document it replaces was wrong for weeks and nothing said so.* + +**A claim here is counted, not reasoned.** The review of 2026-08-31 found a bundle described as +carrying two images that carries three, a bootstrap described as needing six shapes that uses +four, and ten documents calling themselves `designed` while naming lab-proven code. Each was +produced by describing the system from its design instead of reading it. + +**A phase is done when the lab says so**, and the lab keeps a receipt of when it last ran and +against which commits. A phase marked done here whose assertions have not run is a claim about the +past. + +**What is not proven gets said.** Phase 0 is done and its limitation is written into it. A list +that records only progress becomes a list nobody believes. ## Rules of engagement -These exist so the work can run largely unattended without accumulating the kind of -damage this refactor is meant to remove. +Unchanged from the previous version: they were about how work is done rather than what the work +is. ### Autonomous by default -An agent may, without asking: - -- read anything, measure anything, query any database read-only -- create branches, write code and tests, open pull requests -- run the test suite and typechecks -- write and update `hq/` documents +Read anything, measure anything, query read-only. Create branches, write code and tests, run the +suites, and write or update documents here. ### Always stop and ask -- **destroying or overwriting data** — dropping a table, deleting a provision, rotating a - live credential, removing a module from a node +- **destroying or overwriting data** — dropping a table, deleting a provision, rotating a live + credential, removing a module from a node - **merging anything** — every merge is a human checkpoint, without exception -- **a decision the ADRs do not already answer** — record the question in the relevant - research effort rather than picking and moving on -- **any change to `hq/00-META`** — it is stable by nature +- **anything touching a machine outside the lab**, including a configuration change that restarts + something people are using +- **a decision the records do not already answer** — record the question rather than picking and + moving on +- **any change to [`00-META`](../../00-META/)** — it is stable by nature ### Definition of done for every task 1. tests written **and failing first**, then passing 2. typecheck clean in every package the change touches -3. the local mesh (Phase 0) comes up, and the behaviour is demonstrated in it -4. `hq/` updated if the task changed or answered anything documented -5. deployed, and **delivery verified on every node** — not "the pipeline was green" +3. the behaviour demonstrated **in the lab, on real machines** — not asserted +4. documents here updated if the task changed or answered anything recorded +5. delivered, and the **effect** verified — not that a pipeline was green -### Non-negotiables carried from the current system +### Non-negotiables -- **Never edit mesh-managed files on disk.** Use the owning tool. -- **Never write to production databases directly.** Migrations for schema, application - code for data. -- **Every schema change ships twice** — consolidated schema *and* an incremental - migration. -- **Expand, then contract.** Add the new shape, migrate, verify, and only then remove the - old one — never in a single step. -- **A green pipeline proves transport, not effect.** Verify the effect. - ---- - -## Phase 0 — A mesh that runs locally *(prerequisite)* - -Nothing else starts until this exists. Every fault this refactor addresses was found in -production because there was nowhere else to find it. - -| # | task | done when | -|---|---|---| -| 0.1 | Container image for a node runtime | a node process starts in a container and registers | -| 0.2 | Compose topology: broker, registry DB, object store, *n* nodes | `up` yields a mesh that elects a provider node and settles | -| 0.3 | Seed a minimal mesh: nodes, one module, one provision | a module deploys end-to-end with no external service | -| 0.4 | Run the pipeline inside it | a push-equivalent produces a cascade and a deployed artifact | -| 0.5 | Fixtures for the failure modes already known | credential rotation reaching a running session; a provider deploy rotating a shared credential; a migration that ships nothing — each reproducible on demand | - -**Checkpoint:** a human confirms the local mesh reproduces at least one bug from -2026-08-22 before any decomposition begins. - ---- - -## Phase 1 — Make the model expressible - -The decomposition is impossible while a feature is a singleton per module. - -| # | task | done when | -|---|---|---| -| 1.1 | Decision record — named features, per-node opt-in (next free number) | accepted | -| 1.2 | Manifest: declared `features:` with type + directory | a module declares two of one kind and both build | -| 1.3 | Selection: `always` / flavor-selected / `optional` | a node installs a subset; artifacts stay flavor-blind | -| 1.4 | `requires:` moves onto the feature | a schema feature's database is not provisioned where the feature is not installed | -| 1.5 | Assignment carries the opted-in feature set | opting a node in requires no rebuild | - -**Checkpoint:** one existing module converted to declared features, deployed, verified — -before any others follow. - ---- - -## Phase 2 — Draw the boundary the domain already has - -Cheapest first, and each one proves the extraction pattern before the expensive ones. - -| # | task | extracted from | risk | -|---|---|---|---| -| 2.1 | `hal/knowledge` — one store, review workflow ported | hippocampus + noxflow `knowledge_*` | low — additive | -| 2.2 | `hal/stream` — the record; notifications and messaging as views | axon, synapse, notifications, meetings, conversations | medium | -| 2.3 | `hal/agents` — identity, licence, runs, memory, thoughts | noxflow agents, `hal/thoughts` | **high** — touches credentials | -| 2.4 | `hal/work` — what remains of noxflow | noxflow tasks | medium | -| 2.5 | `hal/ai` — provider integration, flavored | `hal/claude*` | medium | - -Each extraction is expand-then-contract: new context alongside, dual-write, verify, cut -over, remove. **Never a move commit.** - -**Checkpoint:** after 2.1, a human confirms the extraction pattern before 2.2 begins. -After 2.3, a human confirms credentials still reach every agent on every node. - ---- - -## Phase 3 — Reclaim the kernel - -Only possible once domains have modules to own their code. - -| # | task | done when | -|---|---|---| -| 3.1 | Move work-domain code out of `hal/sdk` | `workflow-engine.ts`, `task-commands.ts` live in `hal/work` | -| 3.2 | Move provider code out | `claude-credentials.ts` lives in `hal/ai` | -| 3.3 | Move delivery code out | feature handlers, artifact manager, build executor live in `hal/delivery` | -| 3.4 | Decide the residue | ADR: what `hal/sdk` keeps (open question 4) | - -**Measure:** `hal/sdk` line count, tracked per task. Today: **34,636** across **155** -files. - ---- - -## Phase 4 — Separate what the mesh runs from the mesh - -| # | task | done when | -|---|---|---| -| 4.1 | Decide the destination (open question 3) | ADR accepted | -| 4.2 | Cross-repository dependency resolution proven | a catalogue module builds against a published `@hal/*` | -| 4.3 | Move the 91 catalogue modules | this repository contains only mesh contexts | - -**Checkpoint:** move one application first and run it for a week before the rest follow. - ---- - -## Sequencing constraints - -- **0 before everything.** Unverifiable refactors are how this list got long. -- **1 before 2.** Extracting into contexts without per-node features recreates the module - count inside the new names. -- **2 before 3.** A domain can only own its shared code once the domain has a module. -- **2.3 after 2.1 and 2.2.** Agents touch credentials; do it once the pattern is proven on - cheaper contexts. -- **4 last.** It is the only phase that is pure movement, so it is the only one safe to - defer indefinitely. +- **Never edit mesh-managed files on disk.** Use the thing that owns the file. +- **Never write to a production database directly.** Migrations for schema, application code for + data. +- **Every schema change ships twice** — consolidated schema *and* an incremental migration. +- **Expand, then contract.** Add the new shape, migrate, verify, and only then remove the old one. +- **A green pipeline proves transport, not effect.** ## What "done" looks like -Eight contexts. `hal/sdk` holding only what is genuinely cross-cutting. A mesh that stands -up on a laptop. A module count that grows only when the domain does. +The old registry is off. Every module runs on the new mesh, declared rather than scripted. A +machine that fails says what it could not do. And the number of modules grows when the work does, +not when the platform needs somewhere to put something. diff --git a/03-DESIGN/01-to-be/01-end-to-end-testing.md b/03-DESIGN/01-to-be/01-end-to-end-testing.md index e73a48d..967c23a 100644 --- a/03-DESIGN/01-to-be/01-end-to-end-testing.md +++ b/03-DESIGN/01-to-be/01-end-to-end-testing.md @@ -2,11 +2,10 @@ layer: to-be status: in-progress code: [mesh-lab] -updated: 2026-08-23 +updated: 2026-08-31 decisions: - - 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md - - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md - - 02-DECISIONS/0030-the-repository-structure.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0019-how-this-repository-works.md --- # End-to-end testing @@ -35,9 +34,9 @@ today, that is a gap in the vocabulary rather than a reason to privilege that sh ## Two classes of scenario -The design below describes a scenario as a complete mesh — forge, coordinator, delivery cascade -— because what it tests is a module. **That is the larger of two classes, and not the first one -built** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). +The design below describes a scenario as a complete mesh — forge (Gitea), coordinator, +delivery cascade — because what it tests is a module. **That is the larger of two classes, and +not the first one built** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). | | **Bootstrap scenario** | **Full scenario** | |---|---|---| @@ -127,7 +126,7 @@ drifts. a mesh named by the request instead. - **Scenarios must be concurrent and cheap.** Several agents working means several scenarios at once, each needing its own network and nodes. A lab node is a virtual machine - ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)), and snapshots are + ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)), and snapshots are what make repetition cheap — restoring a scenario costs far less than building one. The earlier argument here, that only system containers made this affordable, was superseded: the scale it assumed was invented rather than required. @@ -337,7 +336,10 @@ credential rotation reaches every consumer, that delivery to an absent node is r pending rather than done, that a returning node catches up. These are fewer and change rarely, but they are where the known production faults get encoded so they stay fixed. -The known faults become mesh tests that fail today. That is the Phase 0 checkpoint. +The known faults become mesh tests that fail today. Phase 0 of +[`00-work-breakdown.md`](00-work-breakdown.md) is now complete on this basis — twenty-two +assertions on real machines — and what it does **not** cover is recorded there: no module +from the existing system has run against any of it yet. --- @@ -390,6 +392,38 @@ Everything a node itself does is real, because a node is a real machine. --- +## A suite too expensive to run on every push says when it last ran + +*Written 2026-08-31, from resolving [04-ISSUES/005](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md).* + +This suite needs a machine with a hypervisor. It therefore cannot run on every push, and a suite +that does not run on every push runs **when somebody remembers**. Remembering is not a mechanism, +and the harness this one replaces proves it: it had not built for two and a half months, nothing +said so, and the coverage was assumed rather than checked. + +**The danger is not that the suite breaks. It is that nobody notices it stopped running** — and +that danger belongs to *this* design, not to the harness it retired. + +So three rules, each held by a test: + +**A run leaves a receipt** — when, what passed, what it ran, and the commit each repository was +at. Kept **outside version control**: the question is *has this machine run it*, and a receipt in +git would be a claim about everybody's machine made by whoever committed last. + +**A receipt says why it does not count.** Old, failed, taken against commits the repositories have +moved past, or a run that never raised a machine. Something can be asked, and answers non-zero. +**A receipt that says nothing about something is not a receipt that clears it** — including a +receipt written before it recorded a given fact, which claims nothing rather than everything. + +**The run rebuilds what it tests.** The suite consumes artifacts from other repositories, and an +artifact rebuilt from memory is one rebuilt sometimes. A stale binary reporting success against +rules that have since changed is the same fault wearing different clothes. + +**The general rule, which outlives this suite:** *silence and success must never look alike.* +It is the same rule the host follows about a service that does not exist +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — absence must be distinguishable +from a failure to answer — applied to coverage instead of to a machine. + ## Consequences **Bringing a node into being is part of the framework.** A test creates its own nodes — one diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index 443e161..ce3b7c5 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -2,11 +2,11 @@ layer: to-be status: in-progress code: [mesh-lab] -updated: 2026-08-25 +updated: 2026-08-28 decisions: - - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md - - 02-DECISIONS/0031-the-lab-provides-the-underlay.md - - 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md --- # The scenario declaration @@ -15,7 +15,7 @@ A scenario is a **declaration of an underlay**, plus what to put on it. It is th everything in the lab hangs off, so it is worth getting small. It states what a hosting provider and a home router would provide, and nothing the mesh is -responsible for ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). +responsible for ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). ## Public networks are unrelated, and routed rather than bridged @@ -87,7 +87,7 @@ Four consequences follow, and every one of them shapes this design: address stop corresponding. This is why the mesh dials outward and never inward -([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)), why a hub exists at +([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)), why a hub exists at all, and why a node's endpoint is something a peer **learns** from arriving packets rather than something anyone configures. @@ -258,7 +258,7 @@ otherwise explicit declaration, and it exists because NAT has to run somewhere. It is a **container, not a virtual machine** — a router is scenery rather than something under test, so the fidelity argument that makes a node a virtual machine does not reach it -([ADR 0033](../../02-DECISIONS/0033-a-router-is-scenery-not-a-node.md)). What a router must +([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). What a router must reproduce is kernel behaviour, and a container has the same kernel. **`machines[].at`** — segment and addresses, or a **list** of them for a machine on several @@ -299,7 +299,7 @@ belongs to a router it does not control, and asleep. Whether the overlay survives that, re-forms, and is noticed to have changed endpoint is **observed**, never arranged -([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). +([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). ## Why the addresses are load-bearing @@ -320,13 +320,13 @@ it must be. The format should make getting this wrong hard rather than merely documented: a segment without a `gateway:` is a public segment, and an address in it — including a gateway's `address:` — that is not documentation space is a declaration error, refused before anything is raised. That is -[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied to a configuration +[ADR 0010](../../02-DECISIONS/0010-delivery.md) applied to a configuration file: the failure it prevents is silent, so the check has to be loud. ## The same declaration serves both classes The bootstrap and full scenarios differ **only in `place:`** -([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). Everything +([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). Everything about the underlay is identical, which is what makes one a strict subset of the other rather than a fork. @@ -354,7 +354,7 @@ not first. ## What a scenario deliberately cannot say - **Overlay addresses, the hub, peer configuration.** Outcomes, not inputs - ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)). + ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). - **What a machine is in mesh terms** — server or workstation, its site, its names. Mesh configuration, established by the mesh. - **A host's capability profile.** Detected, never declared. @@ -543,7 +543,7 @@ cannot yet express. Nothing here mentions overlay addresses, which node is the hub, who peers with whom, any name, or any certificate. Research 004 recorded all of those for this topology, and **a scenario must -not state them** ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)): they +not state them** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)): they are what the mesh does, and a scenario that supplied them would be certifying its own work. The absence is the point. Given the declaration above, whether a hub is elected, whether the diff --git a/03-DESIGN/01-to-be/03-scenario-lifecycle.md b/03-DESIGN/01-to-be/03-scenario-lifecycle.md index 52ada40..7a0ecc7 100644 --- a/03-DESIGN/01-to-be/03-scenario-lifecycle.md +++ b/03-DESIGN/01-to-be/03-scenario-lifecycle.md @@ -2,18 +2,18 @@ layer: to-be status: in-progress code: [mesh-lab] -updated: 2026-08-25 +updated: 2026-08-28 decisions: - - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md - - 02-DECISIONS/0031-the-lab-provides-the-underlay.md - - 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0016-the-lab.md --- # Scenario lifecycle The first thing the lab must do, and the only thing it must do before anything else can be written: **materialise a mesh, return it to a known state, and destroy it** -([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). +([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens to one. @@ -39,7 +39,7 @@ The order is not arbitrary — each step needs the one before it to exist: 1. **Segments.** Isolated links, one per declared segment, belonging to this instance and joined to nothing outside it - ([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). + ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). 2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the translation, forwarding and mapping-expiry the declaration asked for. @@ -60,7 +60,7 @@ habit. ## A failed raise leaves the wreckage A step that fails stops the raise -([ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md)) — and **does not tear +([ADR 0010](../../02-DECISIONS/0010-delivery.md)) — and **does not tear down**. Tearing down on failure destroys the only evidence of what went wrong, which is precisely @@ -100,7 +100,7 @@ made after it, and returning undoes it like any other change. ## Reaching in Everything the lab does to a machine goes through the virtualisation layer, never over IP -([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). `exec` runs a +([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). `exec` runs a command on a machine and returns its output. This has one consequence worth stating plainly: **a reachability question is asked from inside**. diff --git a/03-DESIGN/01-to-be/04-lab-installation.md b/03-DESIGN/01-to-be/04-lab-installation.md index 2d00788..85c4f34 100644 --- a/03-DESIGN/01-to-be/04-lab-installation.md +++ b/03-DESIGN/01-to-be/04-lab-installation.md @@ -1,18 +1,18 @@ --- layer: to-be -status: designed +status: in-progress code: [mesh-lab] updated: 2026-09-11 decisions: - - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md - - 02-DECISIONS/0008-a-failed-step-fails-the-job.md + - 02-DECISIONS/0016-the-lab.md + - 02-DECISIONS/0010-delivery.md --- # Installing the lab on a clean machine The lab has prerequisites — a virtualisation daemon, copy-on-write storage, a pool, an identity permitted to talk to it — and it cannot get them from the mesh, because it is where the mesh is -built ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). +built ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). So the lab needs an install path of its own. This describes it, and the shape it has to take is determined by two failures observed while measuring @@ -56,7 +56,7 @@ and unbounded at worst. **The lab refuses to run degraded.** It does not warn and continue: a warning about a slow inner loop is read once and ignored forever, and the loop stays slow. This is -[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied where the failure is +[ADR 0010](../../02-DECISIONS/0010-delivery.md) applied where the failure is performance rather than an error. ## Two ways the prerequisites arrive diff --git a/03-DESIGN/01-to-be/05-the-node-host.md b/03-DESIGN/01-to-be/05-the-node-host.md index df3ed70..c3c7f1f 100644 --- a/03-DESIGN/01-to-be/05-the-node-host.md +++ b/03-DESIGN/01-to-be/05-the-node-host.md @@ -2,16 +2,19 @@ layer: to-be status: in-progress code: [mesh-host] -updated: 2026-08-26 +updated: 2026-08-31 decisions: - - 02-DECISIONS/0030-the-repository-structure.md - - 02-DECISIONS/0036-a-node-is-a-managed-machine.md - - 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md - - 02-DECISIONS/0038-a-node-joins-by-linking-first.md - - 02-DECISIONS/0039-the-link-is-the-security-boundary.md - - 02-DECISIONS/0008-a-failed-step-fails-the-job.md - - 02-DECISIONS/0041-the-host-depends-on-nothing.md - - 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md + - 02-DECISIONS/0019-how-this-repository-works.md + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0010-delivery.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0005-the-node-host.md --- # The node host @@ -22,11 +25,11 @@ Tier 0. The one thing ever installed by hand, and the only thing that changes a A **statically linked binary that requires nothing to be present** — copy it onto a machine and run it, and that is the whole installation -([ADR 0041](../../02-DECISIONS/0041-the-host-depends-on-nothing.md)). Written in Go, because the +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). Written in Go, because the job is system-level and because the host shares no code with any other tier. A single binary with one job: **apply declared state on this machine** -([ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md)). Overlay +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). Overlay membership, packet filtering, packages, services, containers and filesystems are not six concerns it carries; they are six instances of the one. @@ -58,7 +61,7 @@ returns it. Three properties, each following a recorded decision: **A failed step fails the apply.** Not "logs and continues" -([ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md)). A partial apply that +([ADR 0010](../../02-DECISIONS/0010-delivery.md)). A partial apply that reports success is the mesh's most expensive shape. **Each applier reads back.** Setting a value is not evidence the value took. The firewall is @@ -66,7 +69,7 @@ asked whether the rule loaded; conntrack is asked what timeout it holds. This is §5 as a component requirement rather than a review habit. **What was applied is recorded after it works, never before** -([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)). A failed apply leaves +([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)). A failed apply leaves the machine in whatever state it reached, and nothing must claim otherwise. ### store @@ -75,17 +78,17 @@ Local, and **authoritative while disconnected**. Not a cache of the control plan of what this node has applied and what it currently holds. This is structural rather than convenient: if disconnection is an ordinary situation rather -than an exception ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)), the +than an exception ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)), the store is what makes it ordinary. A laptop shut for a week comes back and reconciles; it does not come back and ask what it is. ### link The node's one connection to the control plane, and its security boundary -([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)). +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). It is the broker connection that already exists -([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)) — outbound, +([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)) — outbound, node-initiated, per-node addressed — carrying **per-node identity instead of a shared credential**. The node owns no password. It owns an identity, and that identity is what it presents. @@ -105,7 +108,7 @@ architecture, a network position. capability is real when it is present, running and working, and the difference is the whole point of detecting it. -The profile is what makes [ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md) +The profile is what makes [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) work: a node is a node, and what varies between them is here rather than in the definition. ### inventory @@ -113,10 +116,24 @@ work: a node is a node, and what varies between them is here rather than in the What this machine *is* — its identity, what it holds, what it has applied. Reported upward over the link; never asked downward. +## What it is not + +- It does not decide anything that needs another node. +- It never queries the mesh database. +- It has no listening surface. +- **It does not manage its own unit.** It manages `service` resources and its own unit is one — + the temptation is obvious and it ends with a host stopping itself half way through an apply, + leaving a machine with nothing running to fix it. The installation owns the host; the host owns + everything else. + +**How it is installed, enrolled, run, upgraded and retired is +[`09-the-node-lifecycle.md`](09-the-node-lifecycle.md)**, in full and in one place. This document +is the component; that one is what happens to it. + ## Where a declaration comes from One behaviour, two sources -([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md)): +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)): | Situation | Source | |---|---| @@ -133,7 +150,7 @@ the mesh, and the full peer set arrives derived. ## What a declaration is -Settled by [ADR 0043](../../02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md). +Settled by [ADR 0005](../../02-DECISIONS/0005-the-node-host.md). **JSON**, because the host has no dependencies to spend and the standard library carries no YAML. **An ordered list of typed resources**, each with a stable identity — the order is stated @@ -163,16 +180,57 @@ what it is. No control plane, no declarations, no network. Verifiable immediatel the first node's path, and it is the claim the skeleton's Move 1 rests on and has never proved: that one host can raise the substrate alone. -The first vocabulary is bounded by something the lab makes unavoidable: **a scenario is a -closed address space**, so a resource that must be fetched cannot be applied there at all. So -stage 2 begins with what needs no network — files, directories, service state — and the types -that need artifacts wait on where those come from, which is open in -[`02-scenario-declaration.md`](02-scenario-declaration.md). +Raising the substrate uses **four** shapes — `package`, `container`, `service`, `action` — +counted from the bundle that exists rather than reasoned about. `file` and `directory` are listed +below because they are the cheapest to be sure of and a substrate that needed them would find them +ready; the current bundle simply does not. **All of them are built:** + +| | | | +|---|---|---| +| `directory`, `file` | **built** | no machine dependency at all | +| `service` | **built** | running/stopped **and** enabled/disabled at boot — a unit started but not enabled stops being true at the next reboot | +| `package` | **built** | present, never upgraded, and **never uninstalled** — the host cannot know what else needs it, so dropping one is *forgotten*, not *removed* | +| `container` | **built** | pinned by digest ([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)); identified by a label carrying a digest of the declaration that made it, because a runtime normalises what it is given and that is indistinguishable from drift | +| `action` | **built** | bundle-only ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)); verify is mandatory and is the idempotency check as well as the read-back | + +**A service says what it must reflect, and that is declared state rather than a command.** +`restart-on` names files whose change means the unit must be restarted — because a running service +does not re-read its configuration, and replacing a file, finding the service already running and +doing nothing leaves a machine behaving the way it did before while every check passes. A *command* +to restart would be an action, and the link may not carry one, so this is the shape that rule +leaves rather than a way around it. + +**It may name a file another module put there**, written `.`. The case that needed it: +a resolver restarting when the mesh rewrites the names, which are computed by the mesh and belong +to its module rather than to the daemon's. Without it the daemon serves the names it started with +for ever — every machine that joined afterwards unreachable by name, and every check passing. An +unqualified name still means *my own*, so the common case reads as it always did. + +**An action's verify is the definition of what the action is for**, and the action's own idea of +being finished must be the same one. *Written 2026-08-31, after this went wrong.* If an action +waits on one test and its verify reads back another, the two can disagree — and then the action +succeeds into a state its own verify rejects. The host says so accurately and uselessly: *the +action ran without error and its own verify still fails.* It is intermittent, it reads as a slow +machine, and the remedy people reach for is a longer timeout, which cannot help. +[04-ISSUES/017](../../04-ISSUES/017-an-action-succeeded-into-a-state-its-verify-rejects/00-report.md) +is that, in the one action the whole bootstrap depends on. + +**The parser enforces the boundary rather than the caller remembering it.** `Parse` refuses an +action and is what the link uses; `ParseTrusted` permits one and is what the bundle uses. The +safe path is the default and the permissive one has to be named. + +**The lab still cannot exercise the last three**, and that is now the only thing in the way: a +scenario is a closed address space, so nothing can be fetched there, and its machines carry no +container runtime. All three were instead verified against a real machine — a container created, +labelled, replaced when its declaration changed, exec'd into and removed; an action that exits +zero and satisfies nothing failing the apply. That is lab-installation work +([`04-lab-installation.md`](04-lab-installation.md)) rather than a constraint on the design, but +until it is done the substrate bootstrap has no end-to-end test. **3 — link and store.** The node connects, receives declarations, and holds what it applied. **4 — enrolment.** The one genuinely new mechanism in -[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md); everything else there +[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md); everything else there is configuration of what already runs. Stage 1 is deliberately the smallest useful thing. The lab currently raises **empty machines** @@ -183,7 +241,7 @@ that, and every later stage is tested by a lab that already works. **The lab is the harness.** A scenario places a host on a machine and asserts what it did — against a real hypervisor, with the boundary never mocked -([ADR 0034](../../02-DECISIONS/0034-a-test-defends-a-decision.md)). +([ADR 0017](../../02-DECISIONS/0017-a-test-defends-a-decision.md)). Each decision above owes a test: @@ -195,6 +253,64 @@ Each decision above owes a test: | 0036 — disconnection is a situation | a node cut off and returned reconciles without being re-adopted | | 0008 — a failed step fails the apply | an apply with a failing step reports failure | +## What a machine says about itself, and what the mesh keeps of it + +*2026-08-31, from finding that half of it was being discarded.* + +**A capability is detected and never assumed** ([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)), +so the only account of what a machine can do is the one the machine gave. That account has two +halves and the mesh was keeping one: + +| | | +|---|---| +| **the yes or no** | gates an assignment — *this machine has no seat, and nothing can be installed to fix that* | +| **the detail** | carries a value — `seat: card1-DP-1`, an architecture, an amount of memory | + +They are **one fact read two ways**: *can this run here* and *what should it be configured as*. A +module that must not be assigned without an OLED panel and one that dims itself differently on one +are reading the same line. Keeping only the first read makes the second unanswerable, and there is +nowhere else to get it — the detector is the only thing that looked. + +**The reason an absent capability is absent goes the same way, and it is the half a person needs +most.** *This machine has no container runtime* is the answer; *docker is not installed* is why. +The first is the mesh's to say and the second is only the machine's. + +### The eight, and what each one is evidence of + +*Written 2026-08-31 from `internal/profile/detectors.go`, because the set was implemented and +enumerated in no document. A vocabulary a module writes against, that exists only in code, is one +nobody can write against without reading the code.* + +| capability | what a detection proves | +|---|---| +| `container-runtime` | a runtime is **running**, not installed | +| `package-manager` | the machine's own package manager works | +| `service-manager` | an init that can be asked for state — including *degraded*, which reports on stdout and exits non-zero | +| `firewall` | a filter this host can write rules into | +| `overlay` | the private network can be joined | +| `graphical-session` | a display server **is running** — state | +| `seat` | hardware where one **could** run — and assignment needs this one, not the row above | +| `privileged` | the host can change the machine | + +**`seat` and `graphical-session` are the pair worth reading twice**, because collapsing them is +the obvious economy and it is wrong in both directions: a machine with a seat and no session can +be given a display server, and a machine with a session running is not thereby able to host a +second one. + +**A detection runs something that only succeeds if the thing is *functioning*, never `--version`.** +A version string proves a binary is on disk, which +[`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) +records as false in the way that matters: the package was installed and the daemon was not +running. + +**Never reported and reported nothing stay different.** One machine has not run the host yet; the +other ran it and can do nothing. Both refuse everything that requires a capability, and the +remedies are not remotely alike. + +*Checked by recording a profile with a present capability carrying a value and an absent one +carrying its reason, and requiring both to survive — and by requiring a machine that never +reported to be distinguishable from one that reported an empty list.* + ## Open - **Whether one host can raise the substrate alone.** Move 1 assumes it. Stage 2 tests it, and @@ -206,7 +322,49 @@ Each decision above owes a test: package, a unit, a container and a dataset *are*. Nothing has measured that surface, and it is the residue of the question [`host-size.md`](../../01-RESEARCH/006-mesh-from-scratch/host-size.md) answered. -- **Rescue.** [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md) suggests it is +- **Rescue.** [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) suggests it is a node whose local state is discarded so the mesh re-derives it, and does not decide it. - **What may expire.** An identity needing refresh to stay valid would make a laptop fail for - being a laptop ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)). + being a laptop ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). + +## What was added to the vocabulary, and why each cost was worth paying + +*Written 2026-08-30. Every addition widens what a compromised control plane can express, so the +count is asserted by a test and a change to it is a decision rather than a convenience.* + +Four shapes raise the substrate. Five more exist because most of what a person installs is not a +service: + +| | why | +|---|---| +| **file**, **directory** | the substrate needs neither, and almost everything else does | +| **user** | a shell, a terminal, a chat client, a desktop are a package plus configuration **in somebody's home**. A mesh with no user owns `/etc` and nothing anybody looks at | +| **archive** | a theme is hundreds of files. Inlining them makes every declaration enormous and rewrites all of them when one changes | +| **network** | a module of several containers has to let them reach each other by name, and doing it with an action would create something nothing could ever remove ([ADR 0029](../../02-DECISIONS/0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)) | + +And `file` gained two fields: `bytes`, because a wallpaper is not a string, and `owner`, because a +file in a home belongs to somebody. + +**`user` also makes a login shell declared state.** `chsh` is a command, the link may not carry +one, and a shell that could only be set by hand is a shell the mesh cannot manage — which is most +of the reason to manage a machine. + +### The refusals that came with them + +- **A file says what is in it exactly once.** `content`, `bytes` and `sealed` are exclusive, so + *what is in this file* is answerable by looking rather than by knowing which field wins. +- **Groups are added, never pruned.** The tool that sets them replaces the set unless told + otherwise, which would silently remove every group that makes a login able to use the machine. + A machine's own groups are not the mesh's to know about. +- **An archive is pinned by digest, checked before a single file is written.** This is the one + place the host reaches out on its own — everywhere else it holds one outbound connection and + fetches nothing — so the only thing making those bytes safe to unpack is that they hash to what + was declared. +- **An entry naming a path outside the archive is refused, not sanitised.** Rewriting it to land + inside would put a file somewhere nobody asked for and report success. The first implementation + quietly relocated it, and a test caught that. +- **Symlinks and device nodes are refused rather than skipped**, or an archive needing one arrives + silently incomplete. + +**A partial host does archives and refuses users**: an archive needs a filesystem and a way to +fetch; a user needs a user database the host is allowed to write. \ No newline at end of file diff --git a/03-DESIGN/01-to-be/06-the-control-plane.md b/03-DESIGN/01-to-be/06-the-control-plane.md new file mode 100644 index 0000000..654acc7 --- /dev/null +++ b/03-DESIGN/01-to-be/06-the-control-plane.md @@ -0,0 +1,240 @@ +--- +layer: to-be +status: in-progress +code: + - mesh-control +updated: 2026-08-31 +decisions: + - 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0006-the-substrate-and-the-control-plane.md + - 02-DECISIONS/0008-a-context-owns-its-store.md + - 02-DECISIONS/0019-how-this-repository-works.md +--- + +# The control plane + +Tier 2. The term appears seventy-nine times across this repository and was defined nowhere, +which is `how-we-build` §5 failing on this repository's own vocabulary. + +This document defines it. It does **not** design the contexts inside it; those are open in +[research 006](../../01-RESEARCH/006-mesh-from-scratch/00-overview.md). + +## The definition + +> **The control plane is everything that needs to know about more than one node.** + +That is the whole test, and it is not arbitrary — it follows from +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md). The host applies and +does not decide *because deciding needs knowledge the machine does not have*. So the line falls +exactly there: + +| Question | Whose | +|---|---| +| write this file, with this content, with this mode | the **host** — one machine | +| which nodes should run the store | the **control plane** — needs every node | +| is this unit running | the **host** — one machine | +| which peers belong in this node's overlay | the **control plane** — needs every node | +| what does this machine have installed | the **host** reports; the control plane **records** | +| has this node been unreachable for a week | the **control plane** — nobody else is watching | + +A useful consequence: **anything a single machine could answer alone is not the control +plane's.** If it needs no second node, putting it here is a mistake, and the tier rule will not +catch it because the dependency direction is still correct. + +## What is inside it + +**Seven contexts and one interface** +([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) — +each one earning its place by the test above rather than by being ours: + +| | | needs to know about more than one node because | +|---|---|---| +| **inventory** | nodes, modules, assignments, versions | that *is* the mesh-wide fact | +| **config** | settings, secrets, and deriving them onto nodes | it derives **onto nodes** | +| **connectivity** | overlay, resolution, exposure, filtering, certificates — **specified in full in [`08-connectivity.md`](08-connectivity.md)** | who peers with whom; which node is reachable | +| **provisioning** | resource grants between modules | consumer and provider may be on different nodes | +| **delivery** | source to artifact to node | it targets nodes | +| **observability** | health, logs, metrics, alerts | *unreachable for a week* is nobody else's to notice | +| **identity** | agents, humans, services, authorisation | credentials follow an agent's node bindings and modality | +| **api** | the one interface every surface speaks to | — it is an interface, not a context | + +**What is deliberately not here.** `work`, `knowledge` and `stream` are **mesh-hosted +applications** — first-party, shipped with everything else, and running on the mesh the way +anything else does. A task does not need to know a node exists, and *being ours does not make +something infrastructure*. `ai` is folded into `config`: a provider licence is an ordinary grant. + +**`record` is an open question rather than an eighth entry.** Contexts integrate through it +([ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md)), which makes it load-bearing, +and [research 006](../../01-RESEARCH/006-mesh-from-scratch/skeleton.md) leaves *where it lives* +unresolved — putting it in the substrate risks recreating the circularity the tier design just +removed. Listing it here would settle by naming what has not been settled by arguing. + +**One of the seven is built.** `inventory` owns a database of that name and holds the node records; +the rest do not exist. What it takes to run any of them — the language, and what must already be +running before it starts — is +[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), which also records where the +build stopped and why: at **identity**, because what a node presents to prove who it is is not +decided anywhere, and a migration is the most expensive place in this system to guess. + +**These are contexts, not services.** They are separate in the sense that matters — each owns +its own store, and they integrate through the record rather than by reading one another +([`how-we-build`](../../00-META/how-we-build.md) §4). They are not separate deployables, and +[research 011](../../01-RESEARCH/011-the-module-graph/00-overview.md) records why that +constraint is load-bearing: a single surface can compose them only while there is one interface +in front of them. + +## Nothing outside a context touches its store + +The question this answers: **can a node write to the registry database?** No — and not "only +through one node", which is the weaker arrangement it might be mistaken for. + +> **No node holds a credential to any control-plane store, for writing or for reading.** + +That is not a new rule here; it is four already taken, and it is worth seeing them together +because each one alone reads like a detail: + +| | | +|---|---| +| [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) | the host never queries the mesh database | +| [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) | a node holds its own identity **and nothing else** — the shared database credential every node carries today is the exposure this exists to remove | +| [ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md) | a context is granted only what it **exclusively** owns: no shared writes, no read-only roles | +| [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) | there is no single mesh database, and nothing reads one | + +### So how does anything get in + +**Over the broker, as a message; the owning context writes.** + +``` +node ──event/report──► broker ──► the context that owns that data ──► its own store +``` + +A node reports what it applied, what it holds, and that it is alive. It **states**; it does not +**write**. The difference is the whole security boundary: a node that can write cannot be +prevented from writing anything, and a node that can only state has its blast radius bounded by +what the message vocabulary can say. + +Reads work the same way in reverse — a node is *told*, in declarations. It never asks. + +### Who actually consumes, and who writes + +**The control plane is the consumer. There is one of it, and the context that owns the data does +the write.** + +``` +node ──► broker ──► the control plane, consuming + ├─ a node reported what it applied ─► inventory writes the registry + ├─ a node reported health ─► observability writes its own store + └─ a grant was requested ─► provisioning writes its own store +``` + +Seven contexts, **one deployable** — they are not separate services, so this is one process +consuming and dispatching internally, not seven consumers racing. Each context then writes only +the store it exclusively owns +([ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md)). + +**One consumer is a property worth having**, not just a consequence of +[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md). The as-is records that +*two consumers accidentally sharing one queue silently split the traffic between them, each +receiving half of what it expects* — which has happened, between a module's daemon and its +capability server. With one consumer that class of fault cannot arise. + +**And the broker is the buffer while the control plane is down.** Nodes go on publishing; +messages queue; the control plane drains them when it returns. That is what makes +[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)'s single control plane +tolerable — an outage delays the mesh's *knowledge* rather than losing it. + +**With one consequence that must be bounded before it is discovered:** a queue with no limit +grows until the broker's disk is full, and the broker is the one component every node depends +on. Queues carrying node reports need a maximum length or a message lifetime, and losing the +oldest health report is obviously right where losing the oldest declaration acknowledgement is +not — so the bound is per queue and is not decided here. + +### On volume, which is the real worry underneath + +**Most high-frequency writes are not registry writes, and that is the first thing to check +before designing for throughput.** The registry is `inventory`'s store: nodes, modules, +assignments, versions. Those change when somebody changes something. + +**Logs, metrics and health checks belong to `observability`**, which owns a different store +([ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md)). Sending them to the registry +would be exactly the shared-schema mistake 0045 exists to stop, arriving through the back door +marked *performance*. + +That leaves one genuine funnel: every context's writes go through the process that owns it, and +[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) says there is one of it. +For a mesh of a handful of machines this is not a scaling problem, and **it should not be solved +by giving nodes database credentials** — that trades a bounded problem for an unbounded one. If +it ever binds, the answers are at the consumer: batch, apply backpressure, or move the highest +volume stream out of a relational store entirely. + +**What observability actually stores its data in is not decided**, and it is the one place where +volume genuinely argues against a relational store. + +## What it is not + +- **Not the thing that changes machines.** It decides; the host applies. It never reaches into a + node except through the host. +- **Not a surface.** Tier 3 is how people and agents reach it. It has one interface; the + surfaces are what speak to that interface. +- **Not the substrate.** It *runs on* tier 1 — PostgreSQL, LavinMQ, an OCI registry + ([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) — and cannot start without + them, which is what makes them a lower tier. +- **Not privileged on a node.** It has no more access to a machine than the declaration + vocabulary allows ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). + +## It is also a consumer + +The property that makes tier 2 unlike the others: **the control plane has requirements of its +own.** It needs a PostgreSQL database, an AMQP virtual host, and a bucket — the same things any +module needs, granted the same way. + +That is the circularity the tiers exist to resolve rather than hide: the control plane cannot +provision its own database, because it is not running yet. So its **store** is raised from the +bundle the host carries, before there is a control plane to ask +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md), +[research 011](../../01-RESEARCH/011-the-module-graph/worked-provider.md)). + +Its virtual host and its bucket are **not** in the bundle — by the time they are wanted there is +a control plane to grant them. Whether the bus must come first is +[open](07-the-substrate.md#open), and it turns on whether these contexts talk to each other over +it. + +## Where it runs + +**On nodes, like anything else.** It is not a place outside the mesh; it is modules the mesh +hosts, assigned to nodes by the same mechanism as everything else. + +**One node runs it, and nothing takes over** +([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). The node is assigned, +never elected — no promotion, no quorum, no split brain. + +That is sound rather than merely cheap, because the design already tolerates the control plane +being absent by construction: a node reconciles from **its own** store +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)) and +never needed to ask anybody to hold the state it was last given. So the control plane being down +is not a new failure mode — it is +[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)'s ordinary disconnected +situation, happening to every node at once. **What is lost is change, not operation.** + +The honest half: this node is a single point of failure, recovery is restore rather than +failover, and **certificate renewal is the clock** — an outage outlasting a renewal window expires +every public name. + +## Open + +- ~~**The contexts themselves.**~~ **Decided** — seven, by + [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md). + What remains open is narrower and named there: **where the record lives**, which research 006 + leaves unresolved because the substrate is the one place it must not go. +- **How far it may be split.** One deployable today. Splitting a context out costs the single + interface a surface depends on + ([research 011](../../01-RESEARCH/011-the-module-graph/00-overview.md)). +- ~~**How many run, and what a node does without one.**~~ **Resolved** by + [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md). What remains is + measurement: nothing reports how long the control plane has been unreachable, or how close a + certificate is to expiry — both needed for restore-not-failover to be a plan rather than a + hope. +- **What the interface is.** One interface is stated; its shape, and whether it is request, + subscription or both, is not + ([research 011](../../01-RESEARCH/011-the-module-graph/worked-provider.md)). diff --git a/03-DESIGN/01-to-be/07-the-substrate.md b/03-DESIGN/01-to-be/07-the-substrate.md new file mode 100644 index 0000000..6be2209 --- /dev/null +++ b/03-DESIGN/01-to-be/07-the-substrate.md @@ -0,0 +1,267 @@ +--- +layer: to-be +status: in-progress +code: + - mesh-host examples/substrate-first-node.lock + - mesh-host internal/apply + - mesh-lab test/integration/mesh.test.ts (a bare machine becomes a mesh) +updated: 2026-08-31 +decisions: + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0006-the-substrate-and-the-control-plane.md + - 02-DECISIONS/0007-connectivity.md + - 02-DECISIONS/0008-a-context-owns-its-store.md + - 02-DECISIONS/0019-how-this-repository-works.md +--- + +# The substrate + +Tier 1. Defined the same way [the control plane](06-the-control-plane.md) is, because the same +gap applied: the word was load-bearing and unpinned. + +## The definition + +> **The substrate is what the control plane consumes and cannot grant itself.** + +Every module that needs a database asks the control plane's provisioning for one. The control +plane needs a database too — and it cannot ask itself, because it is not running yet. That +circularity is not an awkwardness to work around; it *is* the definition. Anything on the wrong +side of it must be raised some other way, and the other way is the bundle the host carries +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). + +The test, applied: + +| | control plane needs it | can it grant itself one? | | +|---|---|---|---| +| a relational store — **PostgreSQL** | its own state lives there | no — provisioning needs the store | **substrate** | +| a message bus — **LavinMQ** | it reaches nodes over it ([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)) | no — it cannot grant itself a virtual host | **substrate** | +| ~~an object store~~ | ~~artifacts and blobs it delivers~~ | — | **not substrate** — [ADR 0028](../../02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) | +| ~~an image registry~~ | ~~images it delivers to nodes~~ | — | **not substrate** — needed to operate, not to start ([ADR 0033](../../02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md)) | +| ~~an identity provider~~ | ~~only if it delegates authentication~~ | — | **not substrate** — it delegates to nothing ([ADR 0031](../../02-DECISIONS/0031-the-control-plane-authenticates-nobody.md)) | +| ingress — **Traefik** | not to start; only to be reached by name | — it grants itself one afterwards | **not substrate** ([ADR 0007](../../02-DECISIONS/0007-connectivity.md)) | +| anything else the mesh hosts | no | — | not substrate | + +*The object-store row was wrong, and how it was wrong is worth keeping.* It answered *can it +grant itself one* — no, it cannot grant itself a bucket — while assuming the first column. **The +control plane does not need an object store**: it has no S3 client and never has, and artifacts +reach nodes as content-addressed blobs in the registry. The row was inherited from the system being +replaced, where an object store distributed module tarballs, and was never re-tested against the +definition above it. *Both columns must be answered, and the second is true of almost any service.* + +**An object store is an ordinary module**, required through the module graph by whatever wants one. +A mesh with no workload needing one runs none. + +**The role and the product are both written**, here and everywhere +([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). The role is what the argument +turns on — the test above works on roles, and would give the same answers for a different store. +The product is what actually gets installed and pinned, and a design that names only the role +does not record that the choice was ever made. + +The dependency is on the **protocol**, not the product: AMQP for the bus, S3 for the object +store, the OCI protocol for the registry. That is what keeps the naming safe rather than a +commitment that cannot be revisited — replacing one is a substrate migration, not a redesign. +The store is the exception, and the exception matters: the provisioning model uses databases, +roles and schemas as PostgreSQL means them, so it is the one member that is not a swap. + +## What that resolves + +**Four or five?** [Research 006](../../01-RESEARCH/006-mesh-from-scratch/00-overview.md) asks +whether the identity provider is a substrate service, and the test answers it *conditionally* — +which is the honest answer rather than a number. + +- If the control plane **delegates** authentication, it cannot serve anybody before the provider + exists, and it cannot grant itself a client. **Substrate.** +- If it **authenticates natively**, the provider is an ordinary hosted service like any other. + **Not substrate.** + +So the count follows from a design decision that has not been taken, and the record should say +that rather than assert four. + +**Why not "important infrastructure".** An identity provider, a mail server and an analytics +service are all infrastructure by any ordinary reading, and none of them are substrate — the +control plane starts and runs without them. *Important* is not the test; *the control plane +cannot obtain it* is. + +## What the substrate is not + +- **Not tier 0.** The host raises the substrate; it is not part of it. The host carries the + declaration that brings the substrate up, and depends on nothing. +- **Not the control plane.** These are services with no knowledge of the mesh. A store does not + know what a node is. +- **Not a place for logic.** The skeleton is explicit: tier 1 is *declarations only, no logic of + its own.* A substrate service is an upstream image, pinned, with configuration. +- **Not privileged.** The substrate is provisioned *from* by the control plane and grants + nothing on its own initiative. +- **Not the mesh's supply of anything** + ([ADR 0028](../../02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md)). + A substrate service and a module of the same product are **different instances**. The mesh's own + PostgreSQL and a PostgreSQL a workload was given are two servers, and a node hosting both runs + two containers — expected, not duplication to be tidied away. + + The substrate is raised from the bundle before any mesh exists, so **it is not in the module + graph**: a workload depending on it would depend on something the graph cannot see, cannot rotate + a credential for, and cannot move. It would also put workload data in the store the control plane + keeps its own state in, where a workload that fills a disk takes down the one thing needed to fix + it. + +## The pinned bundle + +`substrate.lock` holds **what must exist before the control plane runs** — which is a smaller +set than the substrate, and the difference is easy to miss. It is the only place in the mesh +where versions are pinned by hand rather than resolved. + +Being substrate and being in the bundle are two different questions: + +| | is it substrate? | must it precede the control plane? | +|---|---|---| +| PostgreSQL | yes — the control plane's own state lives in it | **yes** — there is nowhere to put that state otherwise | +| LavinMQ | yes — it cannot grant itself a virtual host | **yes** — the control plane reaches a node only over the link, and the link is the broker ([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) | +| the OCI registry | yes — it cannot grant itself a repository | no — the first node fetches upstream ([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) | + +The registry is **substrate by role and ordinary by delivery**: by the time it is wanted there is +a control plane, and it provisions it the way it provisions anything. That keeps the bundle small +enough for the review [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) requires — one +substrate image until +[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) established that the +broker has to precede the control plane, and two since. + +*Corrected 2026-08-31, from counting what the bundle holds rather than reasoning about it.* **It +carries three images, not two** — PostgreSQL, LavinMQ, and the control plane itself, which the +sentence above had overlooked by counting only substrate services. The control plane is what the +substrate exists to start, and it is in the bundle for the same reason they are: there is nothing +to fetch it with yet. It also carries seven actions, a package and a service. + +**Why pinned:** the bundle is applied when no mesh exists, so nothing can resolve a version, ask +a registry, or check a constraint. What the host carries must already be exact. + +**Why references and not payload:** the bundle names images by **digest** and the host fetches +them ([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). A first node is +a real machine with a network; the sealed case is the lab, and the lab places images itself. + +Reproducibility comes from pinning the identity of a thing rather than carrying its bytes, which +is what keeps the bundle small enough for a person to read and check. + +## Raising it + +The order, from [research 011](../../01-RESEARCH/011-the-module-graph/worked-provider.md): + +``` +0 a container runtime exists detected — docker or podman — or installed +1 PostgreSQL runs pulled by digest, from the bundle +2 a database per context is created an action, run locally — one today, `inventory` +3 each context's schema is applied an action, against its own database +4 LavinMQ runs pulled by digest, from the bundle +5 a virtual host, a credential, and actions, run locally + a self-signed certificate +6 the control plane starts and only now is there a mesh +7 the registry, and everything else the ordinary path + are provisioned +``` + +**Steps 4 and 5 are why the bundle is not one image** +([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). The control plane cannot +provision the broker, because provisioning means telling a host, and telling a host happens over +the broker. The first node does not escape this by being local: it enrols the ordinary way, by +dialling the broker at the address in its token. + +**Step 2 is one database per context and not one called `mesh`.** A context is granted only what it +exclusively owns ([ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md)), *the mesh +database* names a thing that will not exist +([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)), and a separate +database is a boundary a cross-context join cannot casually cross where a separate schema is not. + +Only PostgreSQL is raised from the bundle, for the reason in *The pinned bundle* above — the +rest of the substrate is wanted only once there is a control plane to provision it. + +**Step 0 is easy to leave out and it is where several things meet.** A substrate service is a +container, so a container runtime must be working before anything else happens — and a runtime +is a *package*, not a container. + +**Which runtime is detected, not chosen** +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)): a machine that +already has one keeps it. On a machine with none, the control plane names the package, because +what it is called differs per system. It is: + +- what the host's capability detection already reports, and the first use of that report by + something other than a person; +- **adopted rather than installed** when the machine already has one with configuration somebody + chose ([research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)); +- a package, which needs the machine's own package manager and a network — both permitted by + [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md). + +So the bootstrap uses four shapes: **package**, **container**, **service** and **action** — +*counted from `substrate-first-node.lock`, which is the only bundle there is*. It had said six, +adding `file` and `directory`, which this bootstrap never asks for. + +All four are built, as are the host's other five +([`05-the-node-host.md`](05-the-node-host.md) stage 2), so nothing in this bootstrap is blocked +on the host any longer — which is the claim that mattered, and it was true either way. + +**Steps 2 and 3 happen before there is a mesh to do them**, which is why provisioning is part of +the bootstrap rather than a service consumers use later. They are **actions** the bundle +declares and the host runs +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)) — so the +host's vocabulary grows by one shape rather than by one resource type per substrate service. + +## Open + +- ~~**Whether identity is the fifth.**~~ **Closed 2026-08-31** by + [ADR 0031](../../02-DECISIONS/0031-the-control-plane-authenticates-nobody.md): the control + plane delegates authentication to nothing, so identity is an ordinary module. With the object + store gone ([ADR 0028](../../02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md)) + the substrate is three, and no member is conditional. +- ~~**Whether the bus must precede the control plane.**~~ **Resolved** by + [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) — it must, and the question as + posed here could not have answered it. This asked whether the control plane's contexts talk to + each other over the bus; they do not, being one process, which under this framing would have + kept LavinMQ out of the bundle. What decides it is how the control plane reaches a *node*, which + is only ever over the link. +- **What issues the broker's certificate at bootstrap.** New, and created by the row above. A + token pins the fingerprint a host must expect before it sends anything + ([`09-the-node-lifecycle.md`](09-the-node-lifecycle.md)), so the broker needs a certificate at a + moment when there is no mesh to issue one and no public name to obtain one for. Self-signed and + pinned is the shape that fits; how it is later replaced by the certificates in + [`08-connectivity.md`](08-connectivity.md) is not decided. +- **How a context added later gets its database.** By then there is a control plane — but one + holding a credential that can create databases holds more than what it exclusively owns + ([ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md)). +- ~~**Whether the host can do step 2.**~~ **Resolved** by + [ADR 0005](../../02-DECISIONS/0005-the-node-host.md). A service + running on this machine is part of this machine, so the scope was never in question — the real + question was whether the host must learn what a database is, and it must not. The bundle + declares an **action**; the host runs it and verifies it, and what a database means stays with + the module that provides one. +- **Whether one host can raise all three.** The claim under stage 2 of + [the node host](05-the-node-host.md), never proved. If it is false, the tier boundary moves. +- **How the substrate is updated once a mesh exists.** Pinned by hand at bootstrap; afterwards + the control plane could deliver it like anything else, and nothing says whether it does. + +## Raised, and observed + +*Written 2026-08-30, the first time a bare machine became a running mesh and something joined it.* + +**It works, and what that means precisely:** a machine with a container runtime and nothing else +applied the bundle its host carries and ended with a store, a database per context, those +contexts' schemas, a broker holding a certificate it generated itself, and the control plane +serving on top of them. Eleven resources, one command, no mesh to ask anything of. + +**Then it joined itself.** The same machine took a token, checked the broker against the +fingerprint pinned in it, generated three keypairs, and enrolled — which is +[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)'s *the first node is a node whose +mesh is not up yet*, observed rather than argued. Its specialness lasted one command. + +**And a credential crossed.** With a second node recorded, the machine was declared the provider +of a database and pushed to over the broker. What arrived and what did not is the whole of the +[secrets argument](../../02-DECISIONS/0009-modules-and-the-graph.md), measured on a real machine: + +| | | +|---|---| +| the password, in plain text | **on the machine only**, one file, mode 0600 | +| in the declaration that crossed the broker | absent | +| in the control plane's database | absent | +| in what the node reported back | absent | + +**One fault, and it was in the joining.** The token did not say what the mesh calls the machine, +so enrolment needed a flag its own help said it did not — and failed at the broker with an empty +username. Recorded in ADR 0004 as the fifth thing a token carries. \ No newline at end of file diff --git a/03-DESIGN/01-to-be/08-connectivity.md b/03-DESIGN/01-to-be/08-connectivity.md new file mode 100644 index 0000000..41e059d --- /dev/null +++ b/03-DESIGN/01-to-be/08-connectivity.md @@ -0,0 +1,574 @@ +--- +layer: to-be +status: in-progress +code: + - mesh-control internal/catalogue/filtering.go + - mesh-control examples/route-proxy + - mesh-control internal/identity/authority.go + - mesh-host internal/identity/serving.go + - mesh-host internal/apply (the service that reflects a rule set) +updated: 2026-08-31 +decisions: + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0007-connectivity.md + - 02-DECISIONS/0007-connectivity.md + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0007-connectivity.md + - 02-DECISIONS/0006-the-substrate-and-the-control-plane.md +--- + +# Connectivity + +One of [the control plane's](06-the-control-plane.md) ten contexts, and the one with the most +moving parts: **overlay, resolution, exposure, filtering, certificates.** + +It is written as a whole because the five are one design. They share inputs, they must agree, and +every one of them today is computed in a different place by a different module from a different +copy of the same facts. + +## Why it is control-plane work + +Apply [the test](06-the-control-plane.md) — *everything that needs to know about more than one +node* — to each responsibility: + +| | needs to know | whose | +|---|---|---| +| **overlay** — who peers with whom, at what address | **every node**, and which of them can be dialled | control plane | +| **resolution** — which name is which node | **every node** | control plane | +| **exposure** — which public name reaches which container | **which node is publicly reachable** ([ADR 0007](../../02-DECISIONS/0007-connectivity.md)) | control plane | +| **filtering** — which port is open, to whom | what is assigned here, and the overlay's shape | control plane decides, host applies | +| **certificates** — who may present which name | which name belongs to which node | control plane | + +**Not one of the five can be answered by a machine on its own.** That is the whole reason this is +a context rather than a set of node-local modules — and it is exactly what the current +arrangement gets wrong, by computing all five on the node from a direct database connection. + +## The shape: decided centrally, delivered as files + +Every one of the five resolves the same way, and it is worth stating once rather than five times: + +> **The connectivity context computes the configuration. It arrives over the link as `file` +> resources. The service reads files and knows nothing about the mesh.** + +This costs **no new host vocabulary**. `file`, `directory`, `service` and `container` already +exist; WireGuard, the resolver and the proxy are all *a container or a package, plus files*. + +It is also what removes the last two upward dependencies. +[Research 006](../../01-RESEARCH/006-mesh-from-scratch/host-size.md) counted exactly two modules +opening a direct connection to the control plane's database — `wireguard` and `traefik` — and +they are the reason every node permanently holds a credential to it +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). Both are connectivity +modules. **Closing this context closes that set.** + +### And they are modules, not a second mechanism beside the module system + +*Written 2026-08-29, from building it. The first version was code beside the module system doing +the module system's job, and the fault it produced is the point of writing this down.* + +**A machine was on the private network because it had an address.** Every node that had been +placed got a peer list, whether or not anybody wanted it there, and there was no way to say a +machine should stay off. That is what "special-cased" cost, and it was invisible until somebody +wanted the exception. + +**What made it look unavoidable:** a peer list cannot be written in a manifest. It is derived from +every other machine, so it differs on each one and changes when any of them changes. So the +manifest says its resources are **computed** — it names something in the control plane that works +them out per node — and it is a module in every other respect: assigned, resolved, configured by +settings, and absent from a machine nobody gave it to. + +**What that made possible immediately** is the arrangement below, which the code has: + +| module | provides | requires | claims | +|---|---|---|---| +| the WireGuard one | a private network, **and the mesh's own addressing** | | *the* private network, one per node | +| the names one | name resolution | the mesh's own addressing | | +| `networking` | | both of the above | | + +**Three rather than one, because WireGuard is one VPN of several.** Naming the module after the +job — `networking` — and putting WireGuard inside it is the retired *flavor* idea wearing a +generic name: the second VPN has nowhere to go. So a module is named for what it *is* and declares +what it *does*, and `networking` is the third row — requirements and no files +([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)). + +**Names left the WireGuard module for their own.** They had been delivered inside it, on the +argument that a machine with peers and no names is half on the network. True, and the wrong place +to fix it — names are identical over a *different* private network, so bundling them made one +module out of two things. They require the mesh's **addressing** rather than a private network in +general, because that is what they are computed from: over a VPN that hands out its own addresses +the mesh has nothing to write, and refusing is what stops a machine being given a hosts file that +means nothing on it. + +**And the claim is not decoration.** Choosing a different VPN still installed WireGuard — dragged +back in by the names, which needed addresses only WireGuard hands out — and nobody was told. +Running two VPNs is not always wrong; being *the* one the mesh runs over is singular. So it is a +claim, and the collision is refused by name. + +**And the proxy's half, which was the other module reaching into the database.** A web application +requiring a reverse proxy has to say *which name, which port*, and there was nowhere to put it — +`requires` says a thing must exist and never said what to do with it. A module now contributes to +a requirement, the control plane collects every contribution on a node, and the provider is given +them as a file at a path it named. It reloads when that file changes, by the same `restart-on` the +private network needed when a peer list changed under a running interface. + +**The proxy's configuration is not written by the mesh.** It is given the facts and turns them +into whatever it runs, which is why swapping Traefik for something else touches nothing that +publishes through it. See [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) for the +other direction — handing a credential *back* — which is the larger half and is not built. + +**What is still not a module, and why that is correct.** The host needs none of this. It has an +address and a route before the mesh exists — that is the machine's own networking — and the +broker's address is carried in the token rather than resolved +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). **The one connection that +carries modules cannot itself be one.** Everything above it can be, and now is. + +## The order it comes up in + +The one thing to get right, because everything else depends on it: + +``` +0 the node has an underlay address the machine's own — DHCP, or a provider gave it one +1 the node dials the mesh OVER THE UNDERLAY, at the address in its token +2 it proves itself, and is proved to the link exists (ADR 0004, ADR 0004) +3 the mesh grants it an identity and an overlay address +4 the overlay comes up peer graph delivered as files +5 names resolve resolver config delivered as files +6 filtering is applied derived from what is assigned here +7 routes and certificates once this node has something to expose +``` + +**Step 1 runs on the underlay and never on the overlay.** This is the circularity that must not +be created: the overlay is configured by the mesh, so a link that required the overlay could +never be established on a new node. The link stays on the underlay permanently — it is +outbound-only and carries its own identity, so it needs nothing the overlay provides. + +**Nothing before step 3 can resolve a mesh name**, which is why the token carries an *address* +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). Today this is +patched with an `/etc/hosts` floor written underneath the resolver; under this design there is +nothing to patch. + +**Step 1 has a precondition this document treated as a fact to record rather than a requirement: +the broker's node must be dialable by every node, at a stable address, and so must the hub** +([ADR 0007](../../02-DECISIONS/0007-connectivity.md)). Across the internet that means publicly +reachable; on one network it does not. A mesh whose nodes are all behind NAT cannot be raised, and +a broker node whose address moves invalidates every token issued for it. + +**Whether the link should later move onto the overlay, with the underlay as fallback, is +[open](../../02-DECISIONS/0007-connectivity.md).** It is a decision rather than a derivation: the +gain is which network carries bytes, not what an attacker can reach, since the link is already +encrypted against a pinned fingerprint. + +## 1 — The overlay + +**What is decided:** the peer graph. For every node: its overlay address, which peers it holds, +which of those it may dial, and which must dial it. + +**Inputs, all declared:** + +- **reachability** — an endpoint, or none + ([ADR 0007](../../02-DECISIONS/0007-connectivity.md)). Not + inferred from the address shape, which is wrong for carrier-grade NAT, wrong for IPv6, and + wrong for a routable address behind a closed firewall. +- **site** — where the machine physically is, or nothing if it roams. +- **role** — hub or not, **declared**. Today it is inferred from an address prefix, which means a + renumbering is an outage and nothing can be asked which node is the hub. + +**Keys.** Each node generates its own keypair. **The private key never leaves the machine**; the +public key is published to the mesh. This is already true and it is already right — it is +[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)'s *a node holds its own +identity* applied to the overlay, and it means the control plane computes a graph it cannot +itself impersonate. + +**Shape: a hub, with direct peering between co-located nodes.** + +| | | +|---|---| +| two nodes at the same site | peer **directly**, host-routed, with a keepalive | +| everything else | routes through the **hub** | +| a node with no site — it roams | **hub only** | + +**Roaming is hub-only deliberately, and the reason is a property of WireGuard rather than a +preference: there is no failover.** A more specific route to a dead endpoint blackholes; it does +not fall back to the general one. So a node whose location changes gets exactly one path, because +two paths would mean one of them silently swallowing traffic. + +**What the host receives:** an interface configuration and a peer list, as files. It does not +compute them, and after this it holds no credential to the mesh's database. + +### Four things the lab found, none of them visible from the mesh's own state + +*Written 2026-08-29, on the first three machines to actually run this.* + +Each looked like a working network from every angle the mesh can see: the graph was right, the +files were right, the services were up, and every node reported success. + +- **A running interface does not re-read its configuration.** A node joins, every existing node's + peer list changes, each file is replaced — and the service is already running, so nothing + reloads it. Every existing node keeps a network that no longer exists. The declaration has to + say the service must *reflect* the file, which is declared state; a command to restart would be + an action, and the link may not carry one + ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). +- **A hub that shares a site with a spoke was emitted twice** — once as a direct peer and once as + the route of last resort. WireGuard takes one entry per public key, so the interface refuses the + file. The ordinary shape of a small mesh, and in none of the tests written before it ran. +- **Two nodes at one site that neither can be dialled must not peer directly.** Nobody opens the + path, and the direct route is more specific than the hub's, so it wins and blackholes. This + document's own warning, arriving in its implementation: *a more specific route to a dead + endpoint blackholes; it does not fall back to the general one.* +- **The container runtime closes the door the overlay needs.** Docker sets the FORWARD policy to + DROP, so a hub with `ip_forward` enabled still carries nothing between its spokes. The substrate + at tier 1 silently breaks the network at tier 2, and nothing in either tier's state says so. The + hub inserts its own rule above those chains and removes it on the way down. + +**The pattern in all four:** the mesh's picture of the network was correct and the network did not +work. That is the argument for the lab in one line — none of these is reachable by reasoning, and +each was found within minutes of a real machine trying it. + +## 2 — Resolution + +**Two name spaces, and they do not mix:** + +| | resolves to | certified by | +|---|---|---| +| **internal names** | overlay addresses | the **mesh CA** | +| **public names** | whatever the outside world must reach | a **public authority** | + +A node's mesh name is its overlay address. Its public name, if it has one, is a separate fact +used by things outside the mesh — and the separation carries two lessons that were learned +expensively enough to be worth restating: + +- **Mesh names are not multicast names.** A name resolved by local multicast discovery introduces + a delay and a failure mode that appears on one node and not others — the worst shape a fault + can have. +- **A node must not pin its own public name locally.** The duplicate record breaks resolution of + that name for everything else that needs it. + +**What the host receives:** the resolver's configuration, as files, listing every peer's internal +name and overlay address. + +**What goes away:** the `/etc/hosts` floor. It exists because a node had to reach the mesh +database before its own DNS existed; with [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) +nothing needs a name before the link, and a fallback nothing needs is a path nothing tests. + +### Names, and what a container can see + +*2026-08-31, from a container that could not resolve a name every machine could.* + +Internal names are `.internal` — the suffix is the one IANA reserved in 2024, so a name that +leaks into a public resolver fails rather than reaching a stranger's machine. They are computed +centrally, because a name set needs every node at once, and written to each machine's hosts file. + +**A file rather than a resolver**, and the reasoning holds: it works on every Linux, needs no +package, and has no failure mode of its own. The stated trigger for a daemon was *names that are +not one-per-node* — service names, wildcards. + +**But a container does not inherit the machine's names.** It gets its own hosts file holding only +its own hostname. So every name the mesh wrote was invisible to the majority of things that need +one — and *on the machine it always worked*, which is exactly what made it easy to miss. It was +found by a database client on one node failing to resolve another node, on a mesh where both names +were correct and present on both machines. + +**So the mesh gives its names to the containers it declares**, written into each container's own +hosts file by the runtime. That extends the file decision rather than overturning it. Given by the +mesh and not chosen by a module: a module that listed the machines would go stale the day one +joins, and a module that did not would be one whose containers cannot reach anything by name. + +**The boundary, which is deliberate and worth stating:** *declared* containers. A container +somebody starts by hand is not the mesh's to configure, and reaching into every container on a +machine — declared or not — is what a nameserver in `resolv.conf` would be for. + +### The resolver, built + +*2026-08-31.* **A service is reached at `..internal`** — the first label is the +service, the rest is the node — so what resolves is *anything under a node's name*, going to that +node. What routes it once it arrives is a proxy's, and stays separate. + +**The mesh writes the data and runs no daemon.** One wildcard per machine, from the same set that +writes the hosts file. A resolver is third-party software and runs *on* the mesh rather than being +*of* it: the mesh has no business shipping one, choosing which one, or knowing its configuration +language. Swapping dnsmasq for unbound changes that module and nothing in the control plane. + +**Two roles, two claims, because they are different things.** systemd-resolved cannot answer a +wildcard at all — it routes the mesh's suffix to something that can. Treating serving and asking +as one role produces a module that cannot work. + +| | claims | | +|---|---|---| +| serving | `the-dns-port` | answers the wildcards | +| asking | `the-resolver-configuration` | decides what the machine asks | + +So *which* resolver is not a mesh-wide decision. One machine can use what systemd already owns and +another can run dnsmasq, and two of either on one machine is refused rather than fought over. + +**Two things a resolver must not do**, both found by a machine rather than by reasoning: + +- **Take an address something else holds.** systemd-resolved holds `127.0.0.53` *and* `127.0.0.54`. +- **Read `resolv.conf` for its upstreams.** Whatever points a machine at the mesh writes the + resolver's own address there, so it becomes its own upstream and every query it cannot answer + loops until its receive queue fills. It needs no upstream: only the mesh's suffix is routed to + it. + +*Checked on two machines, through the path an application takes — nsswitch, files, then DNS — +because the module deciding what the machine asks is half of what is being tested and only that +path goes through it.* + +**That is now the second reason to want a resolver**, and it is a different one from the trigger +above: + +| | | +|---|---| +| names that are not one-per-node | a service named under a machine — `postgres.novox.internal` | +| containers the mesh did not declare | anything a person or another tool starts on a node | + +### Which resolver is not a question the mesh answers + +*Written 2026-08-31, after treating it as open when it had been decided two days earlier.* + +**A resolver takes over `/etc/resolv.conf`, which is a singular resource, so it is a claim** — +[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) lists it in the table beside the seat +and pid 1. Choosing between resolved, dnsmasq and unbound is **assigning a module**, per machine, +and two of them cannot both be assigned there: + +> `resolved-config and dnsmasq both claim "/etc/resolv.conf", and only one thing may hold it per node` + +So there is nothing global to settle and nothing for the mesh to guess. One machine can use what +systemd already owns and another can run dnsmasq, and neither has to know about the other. + +**What the mesh contributes is the part only it can know**: which machines exist and where. That +is `mesh-resolver`, which writes one file and holds no claim, because writing a file takes nothing +over. A daemon module requires that data and claims the resolver — so swapping the daemon changes +that module and nothing else. + +**This was recorded on 2026-08-29 and reopened as an unanswered question on the 31st.** Which is +the argument for the table in ADR 0009 being a table: the pattern is only obvious once seen, and +the cost of not seeing it is inventing a mechanism that already exists. + +## 3 — Exposure + +Settled by [ADR 0007](../../02-DECISIONS/0007-connectivity.md); summarised here because +this is where it belongs. + +**A route is a grant.** A module that must be reachable declares it needs one; the proxy provides +it and hands back the public name. Ordinary +[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) +vocabulary — the mirror of a database grant, where the consumer supplies a target and receives a +name rather than supplying nothing and receiving credentials. + +**A workload on an unreachable node is proxied by a reachable one, across the overlay.** Which is +the case is a mesh-level fact, which is the fourth reason exposure is control-plane work. + +### What was built + +*2026-08-31.* Nothing new in the vocabulary, which was the claim and is now the fact: a route is a +provision, a proxy provides it, and a module that must be reachable requires it. The consumer +contributes the name it wants and the port it listens on; the proxy receives every consumer that +asked; the consumer is told what the provider serves, which is how it knows its own name. + +**One field was missing, and it is the one anything reaching back needs.** A contribution now +carries **where the mesh says that machine is**. A database is reached *by* its consumer, so the +mesh never had to tell a provider where anybody was; a proxy is the other direction — it is told +to send traffic to a consumer and has to open a connection. Without it every provider implementing +a provision would have to know how the mesh names machines, which is a convention leaking into +every module. + +**Exposure and filtering are different questions and a module answers both.** A workload says what +it listens on and who may reach it; separately, it says it wants a route. A module that asked for a +route and not for the port is unreachable by the proxy it just asked for — which the lab +demonstrates, because the machine is already filtering by the time this runs. + +**Withdrawal, which was open above.** The file the proxy is given is the whole truth about who has +a route, so a proxy replaces its table rather than merging. Merging would keep serving a name whose +module was unassigned — and *a stale public name pointing at nothing fails more visibly than a +stale grant* is the reason it must not survive, not a reason to tolerate it. + +**A name a proxy does not serve is refused by saying which it does.** A route that was withdrawn +and a name that never existed are different things, and a bare 404 makes an operator go and read +the mesh to tell them apart. + +*Checked in the lab by a request to the name reaching the workload across the private network and +returning the workload's own answer, then by unassigning the module and requiring the same request +to stop working.* + +## 4 — Filtering + +**Derived from what is assigned here, and from the overlay's shape** — a node's open ports are a +consequence of what runs on it and who must reach it, not an independent declaration to keep in +step by hand. + +**A rule names its source** ([ADR 0007](../../02-DECISIONS/0007-connectivity.md)). +A rule with no source is open, and must say so rather than appear to restrict something. `scope:` +is removed rather than implemented: five manifests carry it today, it is referenced by no code, +and it is the clearest instance in the repository of *an unenforced rule is indistinguishable +from a wrong one, and costs more, because people believe it.* + +**Unknown keys are refused** — the discipline the host's declaration parser already has +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)), and +the one manifests lack. `scope:` survived because nothing rejected it. + +### What was built + +*2026-08-31. Everything above was the intention; this is what exists, and how each part is +checked. [04-ISSUES/003](../../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) is +resolved by it.* + +**A module says what it listens on**, as a port, a protocol and a source — `mesh`, `anywhere`, or +`machine`. The source is required and there is no default, which is the whole of *a rule names its +source*: a manifest that omitted it would read as a restriction and be none. *Checked by a manifest +with a port and no source being refused, and by one naming a source the mesh cannot render being +refused as well — the second is what stops a source becoming a comment.* + +**The set is derived per node**, from every module assigned to it, not from the module asking for +it. Where two modules want the same port, the wider source wins and both are still named, because +removing one of them must not read as a reason to close a port the other needs. *Checked by +rendering a node whose firewall module has no ports of its own and asserting another module's port +is in the result; and by giving one port two modules and one source each, and asserting the +narrower rule disappears while both names survive.* + +**What is not declared is closed.** The rule set drops by default. *Checked by naming the input +chain in the assertion rather than the policy alone — the first version of that test passed while +input accepted everything, because another chain in the same file also said `policy drop`.* + +**From the mesh means the machines the mesh has**, as their addresses on the private network, not +as a subnet. A subnet is a guess that stays wrong quietly; the address set shrinks when a node +leaves and nobody edits anything. A machine that asks for `mesh` where the mesh knows no addresses +is **closed and told so in the file** — widening it would open a port nobody asked to open, and +dropping it silently would close one somebody did. + +**Three things it deliberately does not do**, each of which looked right and would have broken +something: + +| | why not | +|---|---| +| decide what the machine **forwards** | the container runtime writes its own forwarding rules and a second policy is consulted alongside them, so a drop here stops every container on the node — the control plane included. Nothing in a manifest says what a machine routes, so there is nothing to derive it from either | +| **flush the ruleset** when loading | that empties every table on the machine, the runtime's among them. Only the mesh's own table is replaced, and it is declared empty first so the replacement works on a machine loading one for the first time | +| carry a **command to load itself** | the link may not carry an action ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). A service is declared to reflect the file instead, so replacing it restarts what loads it — the shape that rule leaves, used here for the first time for its real purpose | + +**A module that wants a rule set brings the unit that loads it.** Found the hard way: the unit a +distribution packages for nftables runs, applies the rules and exits, so it is neither running nor +stopped — and a host asked for a service that is "running" reports, quite correctly, that it is +stopped. Every packet was filtered exactly as declared and the machine was marked as not doing what +it was told. + +**The vocabulary has no word for "ran, did its job, and exited"**, and that is a real gap rather +than a wording problem: the whole class of configuration-applying units — packet filters, sysctl, +tmpfiles — is shaped that way. Until there is one, a module ships a unit that stays, which is also +the better shape: how a machine enforces rules is a fact about the machine, and the mesh has no +business depending on what a distribution happens to package. + +**One rule is derived from the overlay's shape rather than from what is assigned: a hub's own +listening port.** A hub accepts inbound connections from every node at other sites; a machine that +is not a hub dials out and needs nothing open, because a reply to a flow it started is already +accepted. The two want different rules on an *identical module*, so `listens` — a static field — +cannot say it. The machine a static answer gets wrong is the one facing the public internet, which +is the machine that most needs filtering. + +*Recorded as a gap on 2026-08-31 and closed the same day.* **A computed module now contributes +listens the way it contributes resources.** The port comes from the endpoint, which is where the +interface takes its `ListenPort` from — one source, so a rule set cannot open a port the interface +is not on. It is open to *everywhere* deliberately: a node at another site is not on the private +network until this port lets it on, so restricting it to the mesh would be a rule that can never +be satisfied by the thing it exists for. + +**A generator that cannot say what a machine opens is refused, not read as silence.** Closing a +port on the evidence of a failure to look is how a machine is severed by a fault somewhere else — +and the machine it would sever is the hub, whose only route to being repaired is the network it +just closed. + +*Checked by filtering the hub and then requiring the mesh to keep working: a declaration still +reaches the other machine, and the other machine still reaches the hub. A rule file that looks +right and a mesh that has stopped are exactly what that guards against.* + +**And it is enforced, which is what separates this from `scope:`.** Checked on two real machines: +two ports opened, one declared, and from the other machine the declared one answers and the +undeclared one does not — then the module is removed and the port closes with nobody editing a +rule. *A rule set that is written but never loaded passes every check that reads the file, which +is why the check reads packets.* + +## 5 — Certificates + +**Two authorities, kept separate on purpose.** + +| | issued by | for | +|---|---|---| +| **public names** | a public ACME authority | anything outside the mesh reaches | +| **internal names** | the **mesh CA** | node-to-node, over the overlay | + +**The split is not collapsed, including in the lab.** A single-CA lab would hide any bug living +in the split, so the lab runs its own ACME issuer on its public segment and keeps the mesh CA +unchanged ([research 004](../../01-RESEARCH/004-lab-network/00-overview.md)). + +**Public issuance requires genuine public reachability.** The HTTP-01 challenge must be answered +at the name being certified, so issuance happens through a publicly reachable node regardless of +where the workload runs — the same asymmetry as exposure, for the same reason. + +**The issuer must be configurable.** Today it is not: the proxy sets no `caServer` and therefore +defaults to the public authority's *production* endpoint. Two consequences, and the second is +worse than the lab problem that found it — every certificate experiment on a real node consumes +production issuance quota, and a retry loop can exhaust it for a week. + +**The mesh CA is not a bootstrap concern.** A joining node verifies the control plane against the +fingerprint in its token ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)), +so nothing needs the CA before membership. It certifies internal names afterwards, and that is +all it does. + +### What was built + +*2026-08-31.* + +**A node generates a fourth key**, and the reason is the one the other three already give: *a key +used for two purposes is one rotation away from breaking the other.* The identity key would work +for TLS and reusing it would mean rotating a node's identity every time its certificate is +replaced. The private half never leaves the machine; the mesh is told the public half at +enrolment. + +**So there is no certificate request and nothing to seal.** The mesh signs a statement binding a +public key to a name it alone assigns, which is the whole of what a certificate authority does. +It issues rather than stores: the node's key does not change, so signing again produces an equally +valid certificate and there is nothing to keep in step. + +**A machine with no name inside the mesh is refused**, not given a certificate for nothing. A +certificate for a name nothing resolves is a certificate nothing can check. + +**And the key is stored in the format a server reads** — PKCS#8 PEM, not the host's own encoding. +That is not an implementation detail of whoever writes the file: the file exists *because +something else reads it*, so the format is the interface +([04-ISSUES/014](../../04-ISSUES/014-a-key-that-is-present-and-unusable/00-report.md)). + +*Checked by a real handshake between two machines: one serves on its internal name with the key it +generated, the other verifies against the mesh's authority and nothing else. Every cheaper check +passed while the server could not start — the key was present, the certificate was valid, and +nothing read either the way a server would.* + +## What this removes + +The list is worth having in one place, because it is most of the argument: + +- **The last two direct database connections from nodes** — `wireguard` and `traefik`, the only + two, both connectivity. +- **Therefore the database credential on every node**, and the object-store credential beside it. + [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)'s central claim becomes + true rather than aspirational. +- **The `/etc/hosts` floor**, and the bootstrap circularity it patched. +- **Hub election by address prefix**, and the silent no-hub failure when nobody knew the + convention. +- **The RFC1918 inference**, and the lab substitution that existed to satisfy it. +- **`scope:`**, and the class of manifest key that means nothing. + +## Open + +- ~~**What happens when the hub is down.**~~ **Resolved** by + [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s + matching question — they were one question. Nothing takes over. WireGuard has no failover, the + hub is declared rather than elected, and non-co-located paths stop while co-located direct peers + and every already-assigned workload keep running. The recovery path is restore, and its deadline + is certificate renewal. +- **Renumbering the overlay.** Made *possible* by declaring the hub rather than inferring it from + an address, but no procedure exists, and a graph delivered node by node has an ordering problem + while it is half-applied. +- ~~**Revoking a route** when a module is unassigned.~~ **Resolved** 2026-08-31 — see §3. The file + a proxy is given is the whole truth about who has a route, so a route does not outlive the module + that asked for it. +- **IPv6.** [ADR 0007](../../02-DECISIONS/0007-connectivity.md) makes + it expressible; nothing here says the overlay or the resolver handle it. +- **Reporting declared-versus-observed.** ADR 0007 makes the disagreement detectable and does not + say who looks or what they are told. diff --git a/03-DESIGN/01-to-be/09-the-node-lifecycle.md b/03-DESIGN/01-to-be/09-the-node-lifecycle.md new file mode 100644 index 0000000..ab703ac --- /dev/null +++ b/03-DESIGN/01-to-be/09-the-node-lifecycle.md @@ -0,0 +1,696 @@ +--- +layer: to-be +status: in-progress +code: + - mesh-host internal/link/run.go + - mesh-host internal/link/enrol.go + - mesh-host packaging/nox-mesh-host-resume.service + - mesh-host packaging/nox-mesh-host-network.sh + - mesh-control internal/token + - mesh-control internal/inventory/nodes.go +updated: 2026-08-31 +decisions: + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0010-delivery.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0005-the-node-host.md +--- + +# The node lifecycle + +How a Linux machine becomes a node, stays one, and stops being one. + +[`05-the-node-host.md`](05-the-node-host.md) describes the host as a component. This describes +it as something that runs for years on a machine somebody else also uses — which is where the +questions that were not being asked live. + +## The states + +``` + unmanaged ──install──► hosted ──enrol──► enrolled ⇄ disconnected + ▲ │ + └─────release──────┘ +``` + +| State | Has | Can | +|---|---|---| +| **unmanaged** | nothing of ours | — it is a Linux machine | +| **hosted** | the host, no identity | apply a local file, apply its bundle | +| **enrolled** | identity, link, store | everything; this is *a node* | +| **disconnected** | identity, store, no link | hold its machine in the last state it was told | + +**Only `enrolled` and `disconnected` are nodes**, and they are the same node in two situations +rather than two kinds of thing +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). **`hosted` is not a +node** — it is a machine with a program on it that has not been told which mesh it belongs to. + +There is no state for *the first node*. That is the point of +[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md): the first node walks the +same path, in an unusual order. + +--- + +## unmanaged → hosted: installing + +In the machine's own idiom, because the package manager and the init file are the system's +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)): + +``` +# Alpine — the intended first node +apk add nox-mesh-host +rc-update add nox-mesh-host && rc-service nox-mesh-host start + +# Arch +pacman -S nox-mesh-host +systemctl enable --now nox-mesh-host +``` + +Two lines each, and the init file behind them is four +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)) — it says +*run the launcher at boot* and nothing else, so a third system is transcription rather than a +port. + +Or, where there is no repository to install from: + +``` +curl -fsSL https:///mesh-host---x86_64.tar.gz | tar -xz -C /usr/local/bin +``` + +**The binary is per system as well as per architecture**, because two of its appliers are. + +**The tarball must never acquire a dependency**, because the mesh's own package repository is +hosted on the mesh. Any route that needs the mesh in order to install the thing that joins the +mesh is a circle — unusable on a first node, and unusable by whoever is repairing a mesh that is +down, which is exactly when it is wanted. + +### The unit it installs + +```ini +[Unit] +Description=Novox Mesh node host +After=network-online.target +Wants=network-online.target + +[Service] +ExecStart=/usr/lib/nox-mesh-host/launch +Restart=always +RestartSec=5s +StateDirectory=mesh-host + +[Install] +WantedBy=multi-user.target +``` + +**Two lines of policy, and that is deliberate** +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). The init is +asked to *start this at boot* and *start it again if it exits*, and nothing else. Both are +expressible in OpenRC, runit, s6 and an Android `init.rc`, so porting this file is transcription +rather than design. + +**`Restart=always` and not `on-failure`**: the host restarts onto a new binary by exiting +*cleanly*, so a supervisor that only restarts on failure would leave every upgraded node stopped, +having successfully upgraded. + +**What the init does not do is decide when to give up.** Counting failed starts and rolling back +lives in the launcher, where it can be tested — `OnFailure=` in a unit file can only be read and +hoped for, and it is the one thing that has to work on a machine where nothing else does. + +**The package owns this file. The host never does.** It manages `service` resources, and its own +unit is a service — the temptation is obvious and it ends with a host stopping itself half way +through an apply, leaving a machine with nothing running to fix it. A declaration naming the +host's own unit is **refused**, and that refusal is a test rather than a convention. + +The line to hold: **the installation owns the host; the host owns everything else.** + +At this point the host is running and **doing nothing**. It has no identity, so there is nobody +to link to and nothing to apply. It answers `profile`, `inventory` and `version`, and waits. + +--- + +## hosted → enrolled: the ordinary case + +``` +nox-mesh-host enrol --token +``` + +The token carries **four** things and is carried by a person +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)): the broker's +address, the fingerprint to expect, **the control plane's signing identity**, and the right to +join once. + +**The fourth is the one this document listed three of.** A node connects to the broker and takes +instruction from the control plane behind it, and those are two different identities. Pinning only +the broker would make the control plane's authority *transitive* — a compromised broker could then +forge declarations, which, since the host applies whatever the link delivers, is the whole machine. +So the transport is verified once at connect, and **each declaration is verified by its signature, +every time**. + +What happens, in order: + +1. the host dials the broker at the address in the token, **over the underlay**; +2. it checks the broker's certificate against the pinned fingerprint — *before* sending anything; +3. it presents the one-time secret **and its own public key**, which the mesh records; +4. it reports its `profile` and `inventory` upward; +5. the control plane decides what this machine should be, and sends a declaration; +6. the host applies it, reads back, and reports. + +**Step 4 is the one that is easy to miss and is what makes step 5 possible.** The control plane +cannot decide what a machine should run without knowing what it *can* run — a graphical session, +a container runtime, an architecture. The profile is not a diagnostic; it is the input. + +**The node computes nothing about the mesh.** It needs one peer to reach; the whole overlay is +derived centrally and pushed down +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md), +[`08-connectivity.md`](08-connectivity.md)). + +### The first declaration is the overlay, and nothing else + +**The mesh makes a node reachable before it makes it useful.** Step 5 is not one declaration +carrying everything the node will ever run. It is two, in order: + +``` +first the overlay — this node's address, its keys, its peers, its names +then everything else — packages, containers, services, files +``` + +Three reasons, and the third is the one that matters when something goes wrong: + +- **It is forced.** A node cannot join the overlay before contacting the mesh, because its + address and peer set are *assigned* — it generates a keypair, publishes the public half, and + receives the rest ([`08-connectivity.md`](08-connectivity.md)). So the overlay is the first + thing the mesh can give it, and it should be. +- **It is what [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) already + says:** *a joining node does the minimum to be reachable, and nothing else.* +- **It is the way back in.** Once the overlay is up, the node is reachable over it — by SSH, by + anything. If a later declaration breaks the machine, there is a route to it that does not + depend on the mesh's control path working. **Sending a large first declaration risks a node + that is broken and unreachable at the same time**, and those two failures are much worse + together than separately. + +### Reachable is not the same as having a control surface + +Worth stating plainly, because the two rules read as a contradiction and are not. + +| | | +|---|---| +| **every node reaches every other node** | over the overlay — SSH, services, ordinary traffic. This is the point of having one | +| **every node consumes from the broker** | its own queue, over its own outbound connection ([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)) | +| **nothing dials a node to control it** | the host has no inbound control surface ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) | + +**[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) is about the control +channel, not about network reachability.** What it forbids is a listening thing that accepts +instructions and changes the machine. A node being reachable on the overlay — the whole purpose +of the overlay — is untouched by it, and so is a person opening a shell on it. + +The distinction is *who can tell this machine what to be*: only the control plane, only over the +link the node opened, only in declarations of known shape. + +--- + +## hosted → enrolled: the first node + +The same path, with the mesh built in the middle of it. + +``` +# 1 — raise the substrate and the control plane from the carried bundle +nox-mesh-host reconcile + +# 2 — the control plane now exists, and issues the first token +mesh-control token issue + +# 3 — the machine joins the mesh it just raised +nox-mesh-host enrol --token +``` + +Step 1 is the bootstrap from [`07-the-substrate.md`](07-the-substrate.md): a container runtime, +then PostgreSQL, then the database, then the schema, then the control plane. It needs no identity +because nothing is being asked of anyone — the host is applying a declaration it already +carries, to the machine it is already on. + +**After step 3 the first node is not special in any way**, which is the property `adopt.sh` and +the bootstrap script never had. Its specialness lasted two commands. + +**And enrolment is exercised on node one.** The path every other node depends on is walked +immediately, against a control plane on the same machine, rather than being written and first +used months later on node two. + +--- + +## Two kinds of host + +Everything above assumes a machine with an init that runs the host at boot. Not every machine +has one ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). + +| | **resident** | **episodic** | +|---|---|---| +| examples | Alpine, Arch | Android | +| started by | an init, at boot | whatever the platform allows | +| supervised by | the launcher | nothing — the platform decides when it runs | +| the link | held open | opened while it runs | +| being stopped | shutdown, or a failure | **ordinary** | +| shapes | all six | `file`, `directory`, `action` | +| can be the first node | yes | **no** | + +**An episodic host being killed is disconnection, not failure.** That is +[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) doing the work it was written +for: reachability is state, not class. Everything the design already does for a laptop that +closes — an authoritative local store, reconcile on start, *last heard from* reported without an +alarm — is what an episodic host needs, at a shorter period. + +**It cannot be the first node**, and that is not a limitation to work around. Every step of +raising a substrate is a `package`, a `container` or an `action` against one, and a partial host +refuses the first two. So `mesh-host-android bundle` returns a file that says so rather than an +empty placeholder waiting to be filled in. + +**Two things this changes for anything reading the mesh.** *Last heard from* is a much weaker +signal on an episodic host — a healthy phone looks like a dead server — so a reader has to know +which kind it is looking at. And a declaration may take a long time to land, which makes +[ADR 0010](../../02-DECISIONS/0010-delivery.md)'s separation of +*outstanding* from *failed* load-bearing rather than tidy. + +**Still open:** how an episodic host is started in practice — an APK with a foreground service, +or Termux with its boot addon — and, first, **what an Android node is for.** A device that can +write files and run commands is not a workload host; it is a presence, or somewhere an agent +runs. Building the start mechanism before deciding that would be building it for nobody. + +## Adoption: what happens to what is already there + +Adoption is not a state. It is what the **first apply** does when it is told to own something a +machine already has ([research 012](../../01-RESEARCH/012-the-minimum-viable-node/00-overview.md)). + +A candidate machine is not empty. It has a package manager, probably a container runtime, +configuration somebody chose. [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) +says the host never touches what it did not create — adoption is the deliberate act of taking +ownership of exactly that, so it is a companion to that rule rather than an exception: + +> *never, unless adoption made it the host's* — with adoption **explicit, recorded, and visible +> in what the host says it owns.** + +Three rules, all earned: + +**The original is kept before anything is written.** A one-way door on a working machine is not +an installation. This is a *never* rule, and it earns that from the worst loss in this record — +a tool acting on a path it did not own. + +**On conflict, the machine's configuration wins.** Adoption always completes; the conflict is +flagged and reconciled afterwards. A machine in use keeps working exactly as it did. + +**Adoption produces a briefing**, not just a result: what it found, what it took over, and what +it could not resolve — with each line marked `ok`, `kept`, `unknown` or `failed`, and the overall +outcome **derived** from the worst line rather than stated alongside it. + +--- + +## enrolled: what running actually looks like + +**Changes are pushed, not polled.** A declaration arrives as a message on the link and the host +applies it then. The link is already open and outbound +([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md), +[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — asking it +repeatedly whether anything has changed would be slower to land *and* constant traffic to learn +nothing. + +| Trigger | Kind | | +|---|---|---| +| **a declaration arrives** | **pushed** | the ordinary path — this is how a change lands | +| **start** | event | the machine may have changed while nothing was running | +| **reconnect** | event | declarations may have been missed | +| **every ten minutes** | periodic | **drift, and only drift** | + +**Why the timer cannot be an event.** Drift is change the *mesh did not make* — somebody edited +a managed file, a distribution upgrade replaced a config, a container was stopped by hand. +Nothing will ever publish a message about it, because whatever did it is not part of the mesh. +Only looking finds it. + +So the two periodic things do different jobs and should not be conflated: + +| | direction | answers | +|---|---|---| +| **reconcile timer** | local, looks at the machine | *does this machine still match what it was told?* | +| **heartbeat** | upward, reports to the mesh | *is this node still here, and what is it running?* | + +**The heartbeat is what makes silence mean something.** A node with nothing to do sends nothing; +without a heartbeat that is indistinguishable from a node that stopped. With one, *last heard +from* is a fact beside every node — which is what +[how long disconnected](#how-long-disconnected-and-who-is-told) reports and what +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) exists +because a stuck node cannot send. + +**Rebooting mid-apply is safe by construction.** The store records each resource *after* it +worked ([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)), so a host that +dies half way through comes back, finds the completed ones already matching, and applies the +rest. The rule that exists to stop the host lying about what it did also makes it crash-safe. + +## Updating what the node holds + +An ordinary declaration. Someone assigns a module; the control plane recomputes what that node +should be and sends it; the host applies the difference and removes what is no longer declared. + +**Removal is not symmetric, and the asymmetry is the design:** + +| | on being undeclared | +|---|---| +| file, directory | **removed** | +| container | **removed** — the host created it | +| service | **stopped**; the unit file is not the host's to delete | +| package | **left installed** — *forgotten*, not removed | +| action | **forgotten** — it left nothing the host owns | + +The host removes what it *made* and leaves what it merely *configured*. Uninstalling a container +runtime because a declaration changed would stop every container on the node. + +--- + +## enrolled ⇄ disconnected + +Not a failure. Not degraded. A situation +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). + +A disconnected node **keeps reconciling against its own store**, so it goes on holding its +machine in the last state it was told to hold. A laptop shut for a week comes back and +reconciles; it does not come back and ask what it is. + +What it cannot do: receive new declarations, be granted anything new, or have its certificates +renewed — which is the clock on the whole arrangement +([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). + +**How long it has been disconnected is a fact the mesh must hold**, and nothing holds it today. +Without it, a node running last month's assignments looks exactly like one that is current. + +--- + +## Rescue + +The host is still a command-line tool, and that is what rescue is: + +``` +nox-mesh-host owned # what do you think you own? +nox-mesh-host apply repair.json # apply something by hand, locally +nox-mesh-host profile # what can this machine actually do? +``` + +`apply FILE` accepts actions, because someone who can write that file and run this binary as +root can already do anything it can. The bound in +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) is on what a +**remote** party may push, not on what a person at the machine may do. + +This replaces the three hand-run scripts that exist today — first node, joining, rescue — with +one binary that has always been the same binary. + +--- + +## enrolled → hosted: retiring a node + +Two cases, and they are genuinely different. + +**Graceful.** The control plane sends a final declaration that names nothing. The host removes +what it owns by the table above, reports, and drops its identity. The machine keeps the host +installed and is back to `hosted`. Nothing is left behind that anybody has to remember. + +**The node is gone.** Stolen, dead, or simply unreachable. The mesh cannot tell it anything, and +by [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) it will go on reconciling +its last declaration **forever**. + +That is the honest consequence of making disconnection ordinary, and the answer is not to make +the host expire. It is that **the node holds nothing that outlives revocation**: its identity is +its own, and every grant it holds is a per-node credential at the provider +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md), +[ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md)). Revoking is done at the +database, the broker, the object store — not on the machine. + +So a lost node keeps *running* and stops being able to *reach* anything. That is the best +available outcome and it is worth stating plainly rather than implying the mesh can reach out and +switch a machine off, which it cannot and should not be able to. + +--- + +## Losing the store + +Worth its own section because the failure is quiet. + +If `/var/lib/mesh-host/state.json` is lost — a reinstall, a replaced disk — the host loses +**its record of what it owns**, not its ability to work. It re-enrols, receives the declaration +again, and re-applies it. + +**Without help, what does not come back is removal.** Resources applied under an older +declaration, whose record is gone, become unowned: the host will not touch them, because it +never touches what it did not create. They would sit there, unmanaged, indefinitely. + +**So the mesh keeps a copy of what each node reports it owns**, refreshed on every apply report, +and hands it back on a rebuild — see [Protecting the store](#protecting-the-store). The store +remains locally authoritative for *operating*; the copy exists only for this. + +--- + +## Upgrading the host + +The host is delivered like anything else +([ADR 0010](../../02-DECISIONS/0010-delivery.md)), and this is worth +walking through because tier 0 looks like it should be special and is not. + +``` +push to mesh-host + │ + ├─ build go build → one static binary + ├─ publish packaged, into the mesh's own package repository + └─ deploy each node's declaration now names the new version + │ + └─ pushed to each node; the host applies it on arrival + (a node that is offline gets it on reconnect) +``` + +**Compared with today.** The current pipeline's third silo runs *once per node* and sends each +one a command to install and start. That is where the as-is records a package install that +404ed while the job went green. Here deploy is **one write** — the declaration changes — and the +installing is the host's ordinary work, which reads back before it records anything. + +**The repository is reachable because a declaration made it so.** A `file` resource writes the +package manager's configuration pointing at the mesh's repository; a `package` resource names +the version. Both ordinary shapes, applied by the same host. **No new resource type**, which is +the test of whether this is really uniform. + +### The restart + +``` +1 pacman installs the new binary the running process is untouched — + Unix keeps the running executable's inode +2 the host verifies the new binary runs `nox-mesh-host version`, as a subprocess +3 it finishes the apply and reports never mid-way +4 it exits 0 having finished, not having been stopped +5 the launcher starts it again on the new binary — it supervises the host + rather than exec'ing it (ADR 0005), so this + needs nothing from the init +6 the new host reconciles on start trigger 1, confirming the machine still matches +``` + +**Step 2 is the one to insist on.** A package can install a binary that does not execute here — +wrong architecture, a libc that is not present. Running it once before committing to a restart +turns "the node never came back" into "the apply failed and said why". It is the same read-back +rule the rest of the host already follows, applied to the one resource that is the host. + +**The host never asks the service manager to restart it.** That is the host stopping itself +part-way through an apply. It stops by finishing. + +**A fleet upgrades over an interval, not at an instant**, because each node restarts when its +own apply completes. A node must therefore report the version it is **running**, not the one +installed — otherwise the mesh believes an upgrade landed at step 1. + +**A version that crashes on start rolls itself back** +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). + +What the init starts is not the host but a **launcher**, and the launcher is where the policy +lives: + +``` +init ──► nox-mesh-host-launch ──► nox-mesh-host + ├─ halted? say so and stop; a person has to look + ├─ count this start attempt + ├─ too many, not yet rolled back? roll back, then start + ├─ too many, already rolled back? halt — the machine is the problem + └─ otherwise start the host +``` + +It reinstalls the version recorded in `known-good`, which the host wrote the last time it +completed a reconcile — and the host clears the attempt counter at the same moment, for the same +reason. + +**The launcher rather than the init's own features**, because this is the one thing that must +work on a machine where nothing else does. A shell script with a counter can be run against a +stub package manager and asserted; `OnFailure=` in a unit file can only be read and hoped for. +It also means the init is asked for nothing but *start* and *restart*, which every init can +do. + +**It rolls back once.** If the previous version also fails, the node stops in a failed state +rather than flapping between two binaries. A second failure is a different diagnosis: the +machine is the problem, not the binary. + +**Why this matters more than it looks.** A host that will not start cannot link, and a node that +is not linking looks exactly like a machine somebody switched off — which is the one condition +this design has deliberately decided not to alarm on. Without rollback, a bad release reaches +every node, each one goes quiet, and the mesh reports a fleet of sleeping laptops. + +--- + +## Details that are easy to get wrong + +Each of these has a wrong answer that looks reasonable, which is why they are written down +rather than left to be worked out. + +### Re-enrolling as the same node + +**A token is issued *for* a node record**, and that is where a re-enrolment is decided. + +``` +mesh-control token issue --node workstation # this machine is that node again +mesh-control token issue --new # a machine the mesh has not seen +``` + +The host does not need to know which it is. It presents a token and receives an identity; what +that identity is bound to was decided when the token was made. + +**Issuing a re-enrolment token revokes the previous identity for that node**, and that is not +housekeeping. Two live identities for one node record is the stolen-laptop case with the thief's +credentials still valid — the case +[`retiring a node`](#enrolled--hosted-retiring-a-node) says is answered by revocation. + +### Protecting the store + +**The host reports what it owns, and the mesh keeps the last report.** + +The store stays locally authoritative — a node operates from its own copy and needs nothing to +do so ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). What changes is that +the mesh holds a **copy for recovery**, refreshed on every apply report. + +So a node that loses its state file re-enrols, receives both the declaration *and* the record of +what it previously owned, and can then remove what is no longer declared. The orphans that used +to be permanently stranded are recoverable. + +**This is a backup, never a source.** The host never reads it to decide anything; it is handed +back only on a store rebuild, and a node that disagrees with it wins, because the node is the +one that can see the machine. + +### How long disconnected, and who is told + +**The mesh records last contact per node; the node records time since it last linked.** Both, +because they answer different questions — the mesh's is *have I heard from it*, the node's is +*how stale am I*, and a node reporting the second on reconnect is how a long absence gets +noticed at all. + +**No threshold and no alarm.** A laptop switched off for three weeks is doing nothing wrong, and +a mesh that alerted on it would train people to ignore the alert. It is a **reported fact** — +`last seen 4 days ago` beside every node — and what counts as too long is a judgement for +whoever is looking, not a constant in the design. + +### Whether a failed adoption line blocks + +**Adoption always completes. A node with a `failed` line is a node, and it is not eligible for +assignment until the failure is resolved.** + +*Flags inform, they do not block* holds for **conflicts** — where the mesh chose deliberately and +the machine still works. A **failure** is different in kind: not *we chose* but *we could not*, +and it gets different treatment for that reason. + +The distinction is between **joining** and **being given work**. Refusing to join makes a +machine in use unadoptable, which is the outcome that rule exists to prevent. Placing work on a +machine where something the mesh needed never happened produces a module that is installed and +does not work — [04-ISSUES/007](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) +arriving from the adoption side. + +### What a briefing is + +**A structured document with prose in it**, held in the node's state and reported to the mesh. +It is the first thing a session on a new node reads, which makes it an interface. + +``` +outcome kept derived from the worst line below, never stated separately +node workstation +adopted 2026-08-27T14:02Z + + ok container runtime docker 27.0, adopted; original config kept at + kept storage driver machine has overlay2, the mesh wanted btrfs — machine wins + unknown firewall ruleset could not be parsed + failed package database locked by another process + +what to look at + The storage driver disagreement is preference, not requirement, so nothing is broken. + The package database was locked; nothing was installed. Re-run adoption when it is free. +``` + +**The outcome is computed from the lines**, so a briefing cannot read *fine* while carrying a +failed line. Two independently written fields drift, and that drift is the fault this repository +keeps cataloguing. + +### Where the enrolment token comes from + +**`mesh-control token issue` prints it once**, to the person running it. Single-use, and it +expires whether used or not ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). + +It is carried by hand — read off a screen, pasted into a terminal. That is the design rather than +a gap in it: its authenticity comes from the channel it travelled, which is what +lets a node verify a mesh it has never spoken to +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)). A token emailed, +committed, or dropped in shared storage has lost the only property that makes it worth carrying. + +**On the first node it comes from the control plane that was raised two commands ago**, which is +the same command against a mesh that is one machine old. + +--- + +## Still open + +- ~~**Automatic rollback of a bad host version.**~~ **Resolved** by + [ADR 0005](../../02-DECISIONS/0005-the-node-host.md): a launcher + counts failed starts and rolls back — shipped by the package, not the host binary, because a + binary that will not start cannot recover itself. It rolls back once; a second failure means + the machine is the problem, not the binary. +- **How a previous declaration is retained and chosen**, which is what rollback of anything else + would use ([ADR 0010](../../02-DECISIONS/0010-delivery.md)). +- **A node returning after months** applies a very large jump in one go. Correct, and untested. + +## Going away and coming back + +*2026-08-31, from being asked whether a machine that drops off has to be adopted again.* + +**It does not, and nothing about it expires.** A disconnected node is the same node in a different +situation ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — reachability is state, +not class. The node holds its own identity and the mesh holds the public half; there is no lease, +no timeout, and nothing that lapses while a machine is shut. A laptop closed for a week comes back +and reconnects, and a declaration sent while it was away is waiting for it. + +**The only thing that forces re-enrolment is a machine losing its own key** — a reinstall, an image +re-cloned. That is deliberate: the mesh then reports what it can no longer seal to rather than +delivering blobs the machine cannot open. + +### The gap that was left, and why it mattered + +**A machine that suspends does not know it has been disconnected, and neither does anything else.** +After a resume the socket looks perfectly healthy from inside the process: no error, no close, +because nothing has tried to send anything yet. Heartbeats discover it twenty or thirty seconds +later. + +For those twenty or thirty seconds the node believes it is in the mesh and is not — and *absence +must never be indistinguishable from a failure to answer* is the rule this whole design is built +on. It recovered on its own, which is why this was a quality gap rather than a fault. It was still +the machine waiting to be told something it already knew. + +**So the machine says so.** Waking, and changing network, both rouse the host. + +| | | +|---|---| +| **it ends the current attempt** | not the wait after it — the process is not waiting, it is sitting inside a connection that will never return | +| **by signal, not by anything listening** | a socket for this would be a control surface on every machine, in exchange for saving twenty seconds, and the security argument rests on there not being one | +| **two rouses at once are one** | a machine suspending and resuming repeatedly must not build a backlog of reconnections to work through | +| **the backoff is not reset** | being roused says the machine changed, not that whatever refused the connection has stopped. A laptop woken on a network with no route would otherwise retry at full speed for as long as somebody keeps opening the lid | +| **`down` does not rouse** | the link is already gone, reconnecting will fail, and the backoff exists for exactly that | + +*Checked by holding a link that never returns on its own — which is precisely what a suspended +connection is — rousing it, and requiring the attempt to end and another to begin. And by running +the dispatcher against every event a network manager emits, requiring it to act on the ones that +change where packets go and on no others.* diff --git a/03-DESIGN/01-to-be/10-delivery.md b/03-DESIGN/01-to-be/10-delivery.md new file mode 100644 index 0000000..f1c97e6 --- /dev/null +++ b/03-DESIGN/01-to-be/10-delivery.md @@ -0,0 +1,217 @@ +--- +layer: to-be +status: in-progress +code: + - mesh-control internal/builder + - mesh-control cmd/mesh-control (build, build --behind, push, status) + - mesh-control internal/inventory/builds.go +updated: 2026-08-31 +decisions: + - 02-DECISIONS/0010-delivery.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0010-delivery.md + - 02-DECISIONS/0010-delivery.md + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0009-modules-and-the-graph.md +--- + +# Modules and delivery + +How a change somebody makes becomes a thing running on machines. + +This is the whole of it, current, in one place. Where a decision record is cited it is for the +reasoning behind a choice, not because the answer is somewhere else. + +## A module + +The unit of delivery: assignable to a node, versionable, replaceable on its own. + +**Not a grouping.** There is no `networking` module containing four things — there are four +modules, named individually, with edges between them. Folders assert relationships; edges record +them, and only edges can be queried or kept true automatically. + +**When several modules always change together**, that means they share an *authority* — one place +that decides for all of them. It does not mean they should be one artifact. Connectivity is the +worked example: one context decides the overlay, names, routes, filtering and certificates, and +`wireguard`, the resolver, the proxy and the firewall remain four modules, because they are +deployed to different sets of nodes. + +> **Coherence is a context. Delivery is a module.** + +## The three edges + +A module's relationships to other modules. Two are declared; one is read from the code. + +| edge | means | declared? | satisfied | +|---|---|---|---| +| **presence** | that thing must exist and be reachable here | yes, in the manifest | at provisioning | +| **instantiation** | that thing makes something for me and hands back credentials — a database, a bucket, a route | yes, in the manifest | at provisioning, and again whenever it must be | +| **build** | I was compiled against that artifact | **no — derived from imports** | **at build, once** | + +**Why the build edge is derived and the others are not.** A runtime edge is an *intention* +somebody has about how the mesh should be wired, and only a person can state it. A build edge is +a *fact about code that already exists* — and a declared list of dependencies drifts from the +imports it describes, so the imports are what is read. + +**Why the build edge is a different kind rather than a variant.** It is fixed inside an artifact +rather than negotiated when something runs, and its only remedy is a rebuild. Nothing can +re-provision it. + +## The core library + +One module everything is allowed to depend on, holding **the mesh's own domain**: a module, a +node, an assignment. Those three are what every context talks about and none of them owns. + +The test for whether something belongs: *would this still mean the same thing in a context that +had never heard of the one it came from?* A node would. A pipeline stage would not — that is +delivery's. A grant would not — that is provisioning's. + +**Types ship with the module that owns them**, not here. A consumer needing `inventory`'s types +depends on `inventory` — one narrow, visible edge — rather than everything depending on a hub +where the relationship cannot be seen. A library everything depends on is expensive to change +whether it holds types or code; what makes it expensive is the fan-in. + +**This stays small on its own**, which is the point of choosing a domain rather than a drawer. A +domain model changes when what the mesh *is* changes, which is rare. *Shared code* changes +whenever anybody writes something reusable, which is constantly. + +## Delivery is a comparison, not a pipeline + +The control plane holds two facts and builds the difference: + +``` +what source exists ─┐ + ├─► differ? ─► build ─► judge ─► declare ─► nodes converge +what has been built from it ─┘ +``` + +**A change becomes a build because source is ahead of artifacts.** Not because a message arrived. +An event makes it fast; nothing makes it necessary — so a missed webhook costs latency and cannot +cost correctness. + +That is the same shape the host uses on a machine, one layer up: + +| | reconciles | against | +|---|---|---| +| the control plane | artifacts | source | +| the host | machine state | declarations | + +**There is no pipeline as a state machine.** No stage list something can be omitted from, and no +run to lose. + +### An artifact is current, or it is not + +> An artifact is out of date when **its source moved, or anything it was built against moved**. + +So what is recorded against an artifact is a commit **and the identity of every artifact it was +built against** — its input closure. That is what makes *is this current?* answerable without +building anything, and what makes the rebuild set computable: take the changed module, follow +inbound build edges transitively, and that is what is stale. In order, because the edges are +directed. + +**A shared change is a cascade, and that is inherent.** One change to the core library +invalidates nearly everything. The ordering comes from the graph, not from a hand-written list of +levels. + +### The verdict + +An artifact may not be declared until something has judged it fit. Two tiers, because one gate +would be both slow and unreliable: + +| | judged by | when | +|---|---|---| +| **the module's own tests** | the build | **always** — this is most of it | +| **the lab** | a raised scenario | when an assertion genuinely needs a mesh | + +**A run that failed for environmental reasons is not a verdict.** A machine that would not boot +says nothing about the artifact, and recording it as *unfit* is the same untruth as recording a +dispatch as a deploy. *Outstanding* and *failed* are different results. + +### Declaring, and converging + +Deploy is **one write**: the affected nodes' declarations now name the new artifact. It is not +once per node, and nothing is pushed to a machine. + +Each host applies what it is told, reads back, and reports. A node that is switched off does it +when it wakes. + +**What a delivery result means:** + +``` +meshboard source X · built from X · fit · declared on 5 · applied on 3, 2 outstanding +``` + +Not *the job went green*. **Outstanding is not failure** — a node that has not applied yet is a +fact with a timestamp, and it resolves itself when the node comes back. + +## What this is designed against + +Every property above answers something that has actually gone wrong, recorded in +[`00-as-is/04`](../00-as-is/04-delivery.md): + +| what happened | what prevents it | +|---|---| +| a merge created no pipeline, and nothing said so | a change is found by comparison, not by an event | +| a package install 404'd from every mirror while the job went green | the applier is the reporter, and it reads back | +| a verify stage was built and never scheduled | verification is not a stage that can be left off a list | +| a service was reported started when the command merely returned | *green proves transport, not effect* — so nothing reports transport | +| the build node parked forever while every other node deployed | there is no fan-out to be asymmetric about | + +## What must exist before this can be built + +Not aspirations — things without which the above does not work: + +1. **The module graph, with build edges.** No graph, no rebuild set and no ordering. +2. **A recorded input closure per artifact**, so currency is answerable without building. +3. **Something that notices a reconciler is not converging.** Below. + +## Open + +- **Does a fit artifact declare itself?** Nothing above says who moves the declaration. If it is + automatic, merging to main deploys to production — which may be wanted, and is far too large a + property to acquire by omission. +- **A reconciler that cannot reach its target retries forever.** A failed job stops and names its + step; a loop is silent. Without something that notices *this has been trying for an hour*, this + design reintroduces the fault it removes. **The largest open risk here.** +- **Reproducible builds.** If rebuilding unchanged source against unchanged inputs produced the + same digest, a cascade would stop at the first module whose output did not move. Without them, + one core-library commit redeploys the fleet with no behavioural change. +- **How a module publishes its own types**, which differs per language. +- **How the control plane upgrades itself.** It declares its own new version and the host applies + it — but if the new one is broken, the thing that would fix it is the thing that is broken. The + host has a launcher for exactly this; the control plane has nothing. + +## What "behind" means, and what it used to mean + +*2026-08-31.* + +The risk this record names is losing **did my change go out?** — answerable today by opening a +pipeline, and something has to replace it or the comparison is worse to live with whatever its +other properties. + +**It was answerable only for the machines that broke.** `push --behind` meant *failed or refused*, +so a machine that applied cleanly and whose declaration has since changed was not behind. For every +machine that worked, the answer was silence — and silence meant both *your change is running there* +and *your change has not been sent*, which is the question unanswered rather than answered. + +**So the mesh records a digest of what it last sent each machine.** A digest rather than the +declaration: what a machine should be is recomputable at any moment, and a stored copy would be a +second account of it, able to disagree with the first. What cannot be recomputed is what was +*actually sent*. + +**Recorded after the send.** A digest kept for something that failed to send would make the machine +look current for a declaration it never received — the failure mode this is meant to remove, +arrived at from the other side. + +**Three situations, kept apart**, because they read differently to whoever is looking even where +the remedy is the same push: + +| | | +|---|---| +| **out of date** | it was sent something, and the mesh would now send something else | +| **never told** | nobody has ever asked this machine to be anything | +| **not worked out** | the mesh cannot say what it should be — not reported here at all, because saying "waiting" about it would invent a comparison. `plan` is where that is answered | + +**`status` says it and `push --behind` acts on it**, and both because the alternative is a flag that +knows something the person reading the status does not. diff --git a/03-DESIGN/01-to-be/11-a-board.md b/03-DESIGN/01-to-be/11-a-board.md new file mode 100644 index 0000000..8348e5c --- /dev/null +++ b/03-DESIGN/01-to-be/11-a-board.md @@ -0,0 +1,142 @@ +--- +layer: to-be +status: in-progress +code: + - mesh-control cmd/mesh-control/board.go + - mesh-control cmd/mesh-control/readable.go +updated: 2026-08-31 +decisions: + - 02-DECISIONS/0008-a-context-owns-its-store.md + - 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md +--- + +# A board + +**A place to see the mesh.** Read from a survey of the one that exists, so what is proposed here +is a shorter list than what is there, deliberately. + +## What the existing board does, and what of it belongs here + +Eight sections. Four are about work and workers and are held back with the rest of that domain; +the other four are about the mesh itself. + +| | what it shows | where it stands here | +|---|---|---| +| **the mesh** | every node, what each runs, what each takes from another, module versions | **everything behind it exists** — it is a reader, not a second source | +| **what is wrong** | machines not doing what they were told, machines not answering | **the page nobody had thought to ask for**, and the one a person opens first | +| **builds** | a build, its stages, its log | **everything behind it exists** — every result is kept, failures included | +| **sessions** | model sessions, their usage, and switching between accounts | [ADR 0024](../../02-DECISIONS/0024-model-access-is-a-provision.md) | +| **channels** | nodes messaging each other | the agent layer | + +## The constraint that matters, and it is not a feature + +**The existing board is one service that reads every context's database.** It joins nodes to +provisions to modules to sessions by querying each store directly, because that is the shortest +path to a page that shows all of them at once. + +That is [ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md) violated by the one +component with a reason to violate it, and the cost is not hypothetical — it is the same cost the +shared library has: **a boundary nothing may cross is a boundary that can move; one thing crossing +it is enough to freeze it.** A board that reads the provisioning tables directly is a board that +breaks when provisioning changes its tables, and the change then gets weighed against the board. + +**So a board reads through interfaces and holds nothing.** Everything on the mesh page above is +already answerable by asking the control plane — what nodes exist, what each resolves to, what it +takes from elsewhere, which module came from which commit. A board that asks those questions is a +client. A board that queries `inventory` is a second control plane with a worse contract. + +**It stores nothing of its own.** No cache that can disagree, no table of "what the mesh looked +like last time". If a question is slow to answer, the answer belongs in the context that owns it, +where everything else asking gets it too. + +## Where it is reachable from, which decides everything else about it + +*Written 2026-08-31. It was assumed throughout and stated nowhere, which is the wrong way round +for the most consequential fact about this component.* + +**The board is published on a public name.** Not reachable only over the private network — on the +internet, behind the reverse proxy, like any other published workload. + +**So its login is a perimeter, not defence in depth.** A board on the overlay alone would sit +inside the boundary that [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) already +calls the security boundary, and a login there would guard a room whose door is inside the +building. This one faces everybody. + +**Which makes the identity provider the mesh's outermost gate.** The board is a presentation layer +over the control plane and the control plane's networked surfaces can change the mesh +([ADR 0035](../../02-DECISIONS/0035-one-implementation-several-surfaces.md)), so **whoever that +provider admits can assign modules, from anywhere.** Said flatly because it is easy to arrive at +one reasonable step at a time and then be surprised by. + +Three things follow, and none of them are the board's own design: + +- **Who may log in, and how, is a decision about the mesh** rather than about an application. A + realm that lets somebody in has let them into the mesh. +- **A public name needs a public certificate**, from an authority the world trusts rather than the + mesh's own — which is why that exists at all. +- **The provider failing is not only "the board is down".** Its configuration going wrong in the + other direction — a realm that admits too much — is a mesh-wide exposure with no local symptom. + +**The command line is unaffected and is the reason this is tolerable.** It authenticates through +nothing, needs no network, and answers to the machine's own login — so the mesh remains operable +by somebody standing at it whatever happens to the gate. *That is the property to protect if the +rest of this is ever traded away.* + +## What it is not + +- **Not the way to change things.** Reading is the whole of it to begin with. Every action the + board could offer already exists as a command, and a button that does something no command does + is a second implementation of a decision. +- **Not a dashboard of graphs.** What a person needs from a mesh is *which machine is not doing + what it was told*, and that is a list, not a chart. + +## The three questions, in order + +Written down because the order is the design. A person opens this when something is wrong, and a +page that led with the third would bury the first: + +1. **Is anything broken?** A machine whose last declaration was refused or partly failed. It has + consequences now. +2. **Is anything not answering?** A machine not heard from. It may be new, switched off or + unreachable — **which is not the same as tried and could not**, and collapsing the two sends + somebody to debug a machine that was never sent anything. +3. **Is anything out of date?** A module behind its source, and the machines running the old one. + A plan for later rather than a problem now. + +**Refused and failed stay distinct all the way to the page.** Refused means the machine is exactly +as it was and what is wrong is in what was sent; failed means it is in a state nobody declared and +what is wrong is on the machine. They are fixed in different places, so a page that said "error" +for both would send half its readers to the wrong one. + +## What was built + +*2026-08-31.* + +**One reading, three ways of saying it.** The questions are asked once, by one function, and +answered as a person's `status`, as its JSON, and as this page. Three implementations of *which +machine is not doing what it was told* would be three chances to disagree about it — and the +disagreement would surface as two people looking at two screens arguing about which machine is +broken. + +**It holds nothing and changes nothing.** Every request reads the mesh now. There is no cache to +go stale, no table of what the mesh looked like last time, and no button: every action a board +could offer already exists as a command, and one that did something no command does would be a +second implementation of a decision. + +**It never touches a context's store.** That is the whole constraint above, kept: the board is a +client of the same functions the commands use, so provisioning can change its tables without the +change being weighed against a page. + +**A board that cannot read the mesh says so.** An empty page says *nothing is wrong* in the one +situation where nobody can know that, so the failure is rendered instead — and it says explicitly +that it is a statement about the page rather than about the mesh. + +**A machine's own words are shown, and are not markup.** They are the whole reason the page is +useful — a board that said only *failed* would send a person to ask the thing they opened the +board to avoid asking. They are also the only text on the page that nobody in this repository +wrote, which is why the escaping is a test rather than an assumption. + +*Checked by giving a machine a declaration it cannot apply and requiring the page to name that +machine, say `failed` rather than `error`, and quote what the host said — then by comparing the +page's own JSON against the command's, because two answers to "which machine is broken" would be +worse than either alone.* diff --git a/03-DESIGN/01-to-be/12-a-module-repository.md b/03-DESIGN/01-to-be/12-a-module-repository.md new file mode 100644 index 0000000..127c231 --- /dev/null +++ b/03-DESIGN/01-to-be/12-a-module-repository.md @@ -0,0 +1,347 @@ +--- +layer: to-be +status: in-progress +code: + - mesh-control internal/builder + - mesh-control internal/catalogue/build.go + - mesh-control internal/inventory/secrets.go + - mesh-control cmd/mesh-builder +updated: 2026-08-31 +decisions: + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0010-delivery.md + - 02-DECISIONS/0005-the-node-host.md +--- + +# A module repository, and what builds it + +**Designed from what the mesh needs, not from what came before.** The system this replaces has a +concept of *features* — several independently-deployable units inside one module — and it is +deliberately absent here. + +## Features are unnecessary, and that closes an open prerequisite + +[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) lists *named features +with per-node opt-in* as a prerequisite, on the grounds that without it "every independently +deployable unit inside a context becomes a module again and the count returns." + +**The premise was right and the remedy already exists in another form.** What features were for is +three things the mesh now does separately: + +| features did | what does it here | +|---|---| +| several deployable units in one thing | **several modules**, which is what they are | +| turning one on for one node | **assignment**, which is per node already | +| keeping related things together | **`requires`**, and a module with requirements and no files of its own | + +`networking` is exactly that last row: it ships nothing, requires a private network and name +resolution, and assigning it brings both. So the module count does not return, because the thing +that made it return — *a module is expensive, so put several things in one* — is gone. A module +here is cheap: a manifest and, usually, nothing else. + +## One file at the root + +`module.json`, and a convention somebody can look for beats a setting somebody has to find. It +says what the module is, what it provides and requires, what it claims, what capabilities it +needs, what it puts on a machine — and, if anything must be produced from the source, what to +build. + +## The manifest in the repository is not the manifest the mesh holds + +A resource names an artifact: + +``` +{"id": "dotfiles", "type": "archive", "artifact": "config", "path": "…"} +``` + +and the built manifest names the thing: + +``` +{"id": "dotfiles", "type": "archive", "source": "…/blobs/sha256:…", "digest": "sha256:…"} +``` + +**Two documents on purpose.** A digest is not knowable until something is built, so a repository +carrying one is a repository whose file is wrong the moment anybody edits anything — and the mesh +would be pinning a value nobody could have checked. The built manifest is derived, and the record +of *which commit it was derived from* is what makes "is this current?" answerable without building +it again. + +The word `artifact` never reaches a machine. The host's decoder is strict and would refuse it, at +the worst possible moment. + +## The builder runs on a node + +**Not in the control plane, and this is the same boundary as everywhere else.** Building needs a +container runtime and a working tree; what the control plane may send a machine is bounded by the +declaration language ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)), and *run this build* +is not in it. The alternative — the control plane holding a container socket — would make it the +one component that can do anything on any machine, which is the property the whole design is +arranged to avoid. + +So the builder is a program a machine runs, given work over the broker like anything else, holding +its own credential and nothing more. + +**A build is work, not state**, and that is why it does not travel as a declaration. Everything +else the control plane sends a node is *what you should be*, reconciled forever. A build happens +once and is finished; as a declaration it would either rebuild on every reconcile or carry "and I +already did this" — state about an event rather than about a machine. + +So it has its own queue, and the answer comes back correlated. **One queue**, so several build +machines share the work and each request is done exactly once, which a routing key per machine +would not give. + +**A build machine has its own credential**, and it is not a node's. It may read the build queue +and write to the mesh exchange, and that is all — a node's queue carries that node's declarations, +and a build machine has no business reading them. + +**The answer goes through the exchange, never the default one.** Permission on the default +exchange is granted per *exchange*, not per queue, so anything allowed to use it can publish into +any node's queue. That is the privilege a build machine most obviously should not have. So an +asker binds its own reply queue to the same routing key and filters by correlation; every asker +sees every result, which is the price of the builder never needing that permission. + +Three properties of the builder that are decisions: + +- **a request is acknowledged only once the answer is away.** A builder that dies mid-build then + leaves the work for another machine rather than losing it with nobody ever hearing why +- **one build at a time.** Five at once against one runtime finishes all five slower than it would + have finished the first, and the queue is what shares work between machines +- **a failure is a result.** A build that fails silently is indistinguishable from a builder that + is not running, and those want completely different responses — the same rule the host follows + about a service that does not exist + +### And it is a module the mesh assigns + +*2026-08-31. Written after `builder issue --node`, which is the part that makes the sentence +"holding its own credential" true rather than aspirational.* + +A build machine is a machine that runs the builder, and there is exactly one honest way to say +which machines those are: **assign it**. So the builder is a module like any other — an image, a +container, a working directory, and a claim so a machine does not end up running two. + +The one thing that could not be a module in the ordinary way is the credential. It is not +generated, because the broker has to have been told about it, and it is not written in a manifest, +because a manifest is public and the same file goes to every machine that ever runs it. So the +mesh **creates the account, seals the URL to the machine that will use it, and discards the +plaintext** — the "given, not generated" case above, and its first user. + +Nothing is printed. A credential shown on a terminal is a credential in a scrollback buffer, and +the copy that matters would then exist in two places, one of which nobody is guarding. + +**What this replaces:** a builder started by hand with whatever credential was to hand, which in +practice meant the broker's administrative account. *A program documented as holding its own +credential and given somebody else's is worse than one with no story at all* — the documentation +is what stops anybody checking. + +*Checked in the lab by assigning it and then asking the mesh to build a module: the credential +file arrives readable only by that machine, names the scoped account rather than the broker's own, +and the build completes — which is the only proof the credential authenticates, because a +container that is up holding a credential it cannot use looks identical from outside.* + +**And it is told what to check the broker against**, not only who to connect as. A mesh's broker +presents a certificate of the mesh's own, which is in no public trust store, so a URL alone reaches +only a broker somebody else vouches for — which is no mesh broker at all. The credential carries +the fingerprint beside the URL: **the same two facts a node's token carries** +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)), for the same reason, arriving by +a path other than the thing being trusted. + +**A builder that is a module cannot see the machine's filesystem.** It runs in a container, so a +local path exists for the machine and not for it. That is not a limitation to work around — it is +the arrangement working: a build machine shares the runtime it was given rather than the machine +it sits on. **A module is cloned from the forge over a URL**, and "build this directory" is a +convenience for a builder somebody started by hand. + +### And the loop is closed + +*2026-08-31.* [ADR 0010](../../02-DECISIONS/0010-delivery.md) replaced a pipeline with a comparison +and named the risk: **losing the question "did my change go out?"**. The mesh could already answer +which modules were behind their source — and then a person read that list and retyped each +repository, which is a person being the loop, and the loop is the thing the pipeline was doing +before it was taken away. + +`build --behind` is the other half, and it is the mirror of `push --behind`: the mesh knows what is +stale, so it builds it. The two forms are deliberately not combined — naming a repository and +asking which need building are different requests, and guessing which was meant would sometimes +build something nobody named. + +**One failing does not stop the others**, for the same reason one broken module no longer blocks a +machine's whole declaration: a mesh where one bad repository holds back nine good ones is a mesh +where nobody dares add the tenth. + +**Each is built from its own recorded ref**, not from the commit the mesh happened to notice. +Pinning to that would quietly turn a tracked branch into a pin — a change of meaning nobody asked +for, arrived at by an implementation detail. + +**Building is not delivering, and the two stay separate.** A machine keeps running what it has +until it is told otherwise; the mesh changing its mind is not a machine acting on it, and +collapsing the two is how a mesh comes to report success for something that has not happened. + +*Checked end to end: a commit, a build, a catalogue entry, and a machine that ends up running what +the source says — with both halves that make the answer trustworthy. It is still running the old +one until it is pushed, and it stops being reported as behind once it has caught up, because a +status that says "behind" for ever is one nobody reads.* + +## What is kept + +**Every result, including the failures.** A failed build that leaves no trace is indistinguishable +from one nobody asked for, and the difference is the whole of whether somebody should be looking +at something. A build that failed before it knew what it was building keeps the repository, which +is what a person goes and looks at. + +Recording is idempotent on the correlation, because a result arrives twice — once as the answer to +whoever asked and once on the exchange, where the control plane is also listening. Two rows would +show one build as two, and which is real is not answerable afterwards. + +That is what a builds view reads, and until it existed there was nothing to read: a result was +answered to the asker and kept nowhere. + +## Three properties that are decisions + +- **A fresh clone every time.** A build reusing a working tree can succeed because of something a + previous build left behind, and that is a build nobody can reproduce. +- **Archives are packed deterministically** — sorted, and carrying no timestamps, ownership or + original paths. Two builds of one commit must produce one digest, or nothing downstream can tell + *this changed* from *this was built again*, and every rebuild looks like a change to every + machine holding it. +- **Nothing is published until everything is built.** Half a module in the store, under a digest + the mesh never records, is reachable, unreferenced, and indistinguishable from something in use. + +## What a module may build, and what it may only borrow + +| kind | is | +|---|---| +| **image** | built from a Dockerfile in this repository | +| **archive** | a directory in this repository, packed | +| **upstream** | an image somebody else built, mirrored into the mesh's own registry | + +**The third exists because a module usually runs software it did not write.** A database module +ships configuration and a provisioner and does not build a database. Naming the upstream reference +directly would need every machine to reach a public registry, and would pin to a tag its owner can +move — which is what pinning exists to prevent +([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). Mirroring is what the +bootstrap already does by hand; this makes it something a module can say. + +An upstream reference with **no tag or digest is refused**: what gets mirrored would be whatever +`latest` means today, and a module pinned to that is not pinned. + +## A module's own secret + +A database has a superuser password, a broker an administrator, a registry an account. **None of +them is *for* anybody** — they are not the credential a consumer is given, and the mechanism that +hands those out has a consumer in the middle of it. + +So a module says what it needs and where to put it — `own-secrets`, keyed by a name of the +module's choosing — and the mesh generates one **per node**, +seals it to that machine and reads it no more than it reads any other secret. Per node +deliberately: a module running on three machines has three passwords, where one in the manifest +would put the same secret on every machine that ever runs it, in a file anybody can read, for ever. + +Made once and kept, or a running database would be handed a password it was not started with. +Remade when the machine's sealing key changes. **Declared and not made is refused**, because a +module whose own credential is silently absent starts, fails to authenticate, and the reason is +three layers from the machine reporting it. + +### It is named for whose it is, not how secret it is + +*2026-08-31, from an audit asking whether the manifest format was becoming hard to hold in the +head.* + +The field was called `needs`, beside `secrets` — which is where a **provision's** credential lands +on a consumer. Both were name-to-path, both held something secret, and the names distinguished +them not at all. **Reaching for the wrong one parsed cleanly and failed somewhere else entirely**, +which is the shape of fault this whole design exists to prevent, sitting in the manifest format. + +The axis that separates them is not how secret they are — both are — but **whose**: + +| | keyed by | belongs to | +|---|---|---| +| `secrets` | the provision it is for | the relationship with another machine | +| `own-secrets` | a name the module chose | this module, and nobody else | + +**A manifest using the old name is told the new one** rather than refused with "unknown field". +Whoever wrote it knew what they meant, and the mesh knows what it is called now. + +### Some of them the mesh cannot make + +*2026-08-31, from making the builder a module — the first thing to hold one.* + +A generated secret is the mesh's, and remaking it costs nothing: **nothing else ever knew the old +one.** That is the assumption the paragraph above rests on, and it is not true of every secret a +module needs. + +A broker account's password exists because **the broker was told about it**. A licence key exists +because somebody bought it. The mesh's job with these is to carry the value to the machine that +will use it and then be unable to read it — the same sealing, from the other direction: **given, +not generated.** + +Treating the two alike is wrong in exactly one place, and it is the place nobody looks. When a +machine rejoins it has a new sealing key, and everything sealed to the old one is remade. Remaking +a *given* secret puts thirty-two random bytes where a working credential was, and every visible +signal says it worked: the mesh sealed a secret, the machine applied it, the file is there with +the right permissions. What fails is a program authenticating to something else, hours later, +with an error that names neither the mesh nor the secret. + +So **where the value came from is recorded, and a given secret is never regenerated.** A rejoined +machine asking for one is refused, naming the remedy — issue it again — because the remedy is a +command somebody runs and no amount of pushing will produce a password the broker has never heard +of. + +*Checked by taking a given secret, changing the machine's sealing key, and asserting the mesh +refuses rather than answers; and by asserting that two ordinary pushes hand back the same value, +without which the refusal would be a secret that never survives at all.* + +## What one assignment gets you + +A database module, written to see whether it could be: + +``` +directory /var/lib/mesh/postgres +directory /var/lib/mesh/postgres/grants +container the database pinned by digest, mirrored +container the provisioner pinned by digest, mirrored +file the superuser password sealed to this machine +file what its consumers asked for +``` + +**The provisioner watches** rather than being invoked. That is what lets it be a module: run once, +it needs something to run it after every declaration — a timer, or a unit wired to a file. +Watching, it is an ordinary long-running service the host already supervises. It polls rather than +watching the filesystem, because the host writes atomically: the file is replaced, so a watch on +the path stops seeing anything after the first replacement, and a watcher that silently stops +working is worse than a poll. + +Writing it found one thing wrong, and it was the manifest rather than the host: a container +declared `restart-on`, which is a service field, and the host refused it by name. **It is right +to.** A container whose own definition changes is recreated, and a file it mounts is read by the +process inside, which is that image's business. + +## Where artifacts go + +**The registry the bootstrap already pulls from**, for both images and archives. An OCI registry +is a content-addressed blob store that also understands images, and an archive is a +content-addressed blob. + +An object store beside it is the right answer for objects that are *mutable*, need per-reader +access, or are not build output. None of that describes a digest-pinned archive, and running a +second service for one kind of immutable blob is two things to run, two to back up, and two ways +for an artifact to be missing. **Overturnable without touching anything else**: a manifest carries +a URL and a digest, and neither says what served it. + +### And the mesh runs it + +*2026-08-31.* Which registry is a **provision**, mesh-scoped: a build machine requires +`artifact-store` and is told where it is, the same way an application is told where its database +is. Nothing is configured with an address. + +This closes the last thing the mesh depended on and did not run. The registry a bootstrap pulls +from belongs to whoever raised the machine; from the moment the mesh has one of its own, an +artifact's home is somewhere the mesh can move, replace and back up. + +**The chicken and egg is the bootstrap's, resolved the same way.** A registry module is an +`upstream` artifact — mirrored from a registry that already exists into the one being started. The +first copy comes from outside, exactly once, and every copy after it is the mesh's. + +*Checked in the lab by assigning it and then asking for `/v2/` — on the machine, and from a second +machine across the private network, because a mesh-scoped provision that only answers locally is +not one. A container that is running is not a registry that replies, and this project has paid for +that distinction once already.* diff --git a/03-DESIGN/01-to-be/13-credentials-and-their-rotation.md b/03-DESIGN/01-to-be/13-credentials-and-their-rotation.md new file mode 100644 index 0000000..c61b400 --- /dev/null +++ b/03-DESIGN/01-to-be/13-credentials-and-their-rotation.md @@ -0,0 +1,107 @@ +--- +layer: to-be +status: in-progress +code: + - mesh-control internal/inventory/secrets.go + - mesh-control cmd/mesh-control/rotate.go + - mesh-control examples/postgres-provisioner +updated: 2026-09-01 +decisions: + - 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md + - 02-DECISIONS/0009-modules-and-the-graph.md +--- + +# 13 — Credentials, and moving them + +*Written 2026-08-31, when rotation was built. The delivery half was already proven; this is the +half [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) records as +unowned, and it was **measurably false** in the system being replaced.* + +## What went wrong before, precisely + +`provision_ensure`, documented as *NEVER rotates an existing secret*, minted a new password on +every adoption and updated **only the provider's row**. Consumers on three nodes held dead +credentials for two days. Two rows for one provision were written 216 ms apart, so at most one +could have matched the live role. The mesh reported success throughout. + +Three separate faults, and it is worth naming them apart because they have different fixes: + +| | | +|---|---| +| **one credential, many holders** | rotating it was necessarily a fan-out, and nothing enumerated who held it | +| **the record moved and the consumers did not** | the change and the delivery were different acts, and only the first happened | +| **nothing said so** | the mesh could not tell a rotated credential from a working one, so nobody looked | + +## What replaces it + +**Every pair has its own credential.** A provision between one consumer and one provider is one +password, made once and kept. So rotating a credential touches one role and leaves every other +consumer alone — and *who holds this* is a query rather than an assumption. That alone removes the +first fault: there is no shared secret to fan out. + +**A consumer is a module on a machine, not a machine.** This was written as though a pair were two +machines, and built that way, and it was wrong in a way that only shows on a real node +([`022`](../../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md)): +a machine running several services against one database server had one credential between them. +The provider refused to plan at all, and the consuming node did not refuse — it gave the first +module a credential and the rest nothing. + +Two modules on one node are as separate as two on different nodes. They are different containers, +with different data, and one login opening both is the thing this page exists to prevent. It is +also what makes withdrawal possible: one role per machine cannot express *this module no longer +has a login and the others still do*. + +**The change and the delivery are one command.** `rotate` discards the credential and sends both +ends, and it does the sending itself. Leaving that to whoever remembers is the second fault +exactly, and the interval in which it goes wrong is unbounded — two days, in the recorded case. + +**It is all-or-nothing.** If any affected machine cannot be resolved, nothing is sent and the old +credential keeps working. A mesh that has not rotated is far better than one that has +half-rotated, and the difference is that the first is obvious. + +**The window is stated rather than hidden.** A role's password changes on the provider and the file +changes on the consumer, and those cannot be simultaneous. So there is an interval in which a +consumer cannot authenticate, and the honest thing is to make it as short as the broker allows and +to say it exists. `status` names who is still behind. + +## The provider makes it true, and the mesh cannot + +The mesh generated the password, sealed it to the machine that must accept it, and **discarded the +plaintext** — so it cannot tell a database to start accepting it. Something on that machine reads +what the host wrote and makes it true. + +That something is part of the module, not part of the control plane. **The control plane decides +and never touches a machine; a provisioner runs on the machine and touches it.** Two files, because +the mesh could not compose a document containing a value it does not have: + +| | | +|---|---| +| a manifest | every consumer, what it asked for, and where its credential is | +| one file per consumer | that consumer's password, alone | + +**Both, or neither works.** A provisioner given the passwords and not the manifest finds a +directory of unexplained secrets and reports that nothing has been granted — which is true, and +reads exactly like a credential that was never delivered. That has now happened once, here. + +**The password a provisioner uses is itself a file the mesh wrote.** Passing it through the +environment needs a person in the middle of the one path that exists so there is not one, and puts +a superuser password where `docker inspect` prints it. + +## How it is checked + +Not by comparing two files. **Two ends holding a matching string proves they agree, not that either +is right** — the recorded fault produced two ends that agreed with each other and not with the +database. + +So the check is three logins against a real PostgreSQL, from the consumer's own machine, over the +private network: + +1. the delivered credential authenticates +2. after rotation, the new one authenticates +3. **the one that was rotated away does not** + +The third is what makes it a rotation rather than an addition. Without it the check passes against +a provider that added a password and removed nothing. + +**Not over loopback.** `pg_hba` trusts anything there, so every password looks correct — a +deliberately wrong one returned a row for an afternoon before that was noticed. diff --git a/03-DESIGN/01-to-be/14-model-access.md b/03-DESIGN/01-to-be/14-model-access.md new file mode 100644 index 0000000..4009227 --- /dev/null +++ b/03-DESIGN/01-to-be/14-model-access.md @@ -0,0 +1,184 @@ +--- +layer: to-be +status: in-progress +code: + - mesh-control internal/licences + - mesh-control cmd/mesh-control/licence.go +updated: 2026-09-05 +decisions: + - 02-DECISIONS/0024-model-access-is-a-provision.md + - 02-DECISIONS/0009-modules-and-the-graph.md +--- + +# 14 — Model access + +*[ADR 0024](../../02-DECISIONS/0024-model-access-is-a-provision.md) decided it and listed four +things the mesh did not have. Written 2026-08-31, when two of them were built. **The other two are +still gaps and are still written as gaps** — the record's own warning is that pretending otherwise +is how a plan becomes a surprise.* + +## What was built + +**A licence is a record, and the first provision no machine answers.** Everything else the mesh +brokers is answered by something running on a node. A hosted model is on nobody's machine and is +reached over the public internet, so the rule that refuses two ends sharing no private network — +correct everywhere else — must not apply to it. A machine on no private network at all can hold a +licence, and that is not a special case to remember: it falls out of the answer not being a +machine. + +**The name is the operator's.** *The personal account*, *the organisation's account*. Those are +names a person uses, and the mesh uses them too, because the whole point is saying **which one** a +consumer uses — and an anonymous credential hanging off a provider cannot be said. Many to many, +so deliberately **not a claim**: two machines sharing an account is the ordinary case rather than +a collision. + +**One provision name for all of them.** A module requires `model-access`, never `anthropic`. A +module that named a provider could not be moved onto a model the mesh runs itself without editing +it — and moving it is the point. + +**A model in a machine's own set answers it locally**, and no record is consulted. That is what +makes *the mesh's own model* an ordinary answer rather than a parallel arrangement. + +### Accept: taking a value the mesh did not make + +Every other credential here the mesh generated, sealed to both ends and discarded. An API key +arrives from a person, and the missing verb was *accept*: **take a value, seal it to each holder, +discard the plaintext.** A mesh that kept operator-supplied keys readably is the arrangement this +project measured and rejected. + +**It seals to the holders that exist at that moment**, and this has a consequence that must be +said out loud rather than discovered: + +> A consumer put on a licence *after* the key was supplied has no key, and **the mesh cannot make +> one** — it discarded the only copy. + +So that state is reported at every point a person could meet it: when the consumer is put on the +licence, in `licence list`, and — decisively — **the declaration is refused** rather than written +without the file. A machine that resolves cleanly and receives nothing fails later, somewhere that +names neither the licence nor the mesh. + +**A key is read from a file or standard input, never an argument.** A key on a command line is a +key in shell history and in every process listing taken while it ran. It is never echoed back: +what is stored is unreadable by whoever holds it, the control plane included, and printing it +would put the one copy that matters on a terminal. + +## Refusing is felt, and that is the design working + +ADR 0024 predicted it: *a mesh holding three ways to reach a model refuses every consumer that has +not said which — which is correct and is a great deal of saying-which the first time.* + +It is correct, and correct is not the same as usable. So the refusal names **the candidates and +the exact command**. The difference between a mesh that refuses helpfully and one that merely +refuses is whether anybody can act on it without going and reading something else. + +## Still gaps + +Unchanged from [ADR 0024](../../02-DECISIONS/0024-model-access-is-a-provision.md), and deliberately +not half-built: + +**A consumer that is not a machine.** *This worker uses that licence* is a binding to an agent, not +to a node. What is delivered still lands on a machine; what is **chosen** is chosen per agent, and +the provisions model has no consumer identity other than a node. What exists today is per module +per machine, which is a step toward it and is not it. + +*2026-08-31: this gap now has named consumers rather than hypothetical ones.* +[ADR 0026](../../02-DECISIONS/0026-the-mesh-has-a-session-of-its-own.md) puts two sessions on the +control-plane node — the node's own and the mesh's — each bound in its own right. See +[`15-the-agent-session.md`](15-the-agent-session.md). + +**And for sessions the gap is already closed, which was not obvious.** A binding is per module per +machine, and this document called that *a step toward it and not it* — reasoning that a machine +cannot name an agent. It cannot; but the two sessions are **two modules**, because they are the +same mechanism started in different context roots and a context root is what a module delivers. +So `(node, module)` tells them apart, and asking for a licence per session needed no new consumer +identity. Checked rather than argued: two sessions on one machine hold different licences, each is +given its own key, and releasing one leaves the other. + +**What is still open is the rest of the gap, and it is the harder half.** A *worker* is not one +per machine — many can run on one, from one module — so `(node, module)` cannot name them apart +and this reasoning does not extend to them. That belongs with +[ADR 0003](../../02-DECISIONS/0003-agents-are-persistent-employees.md), which is unbuilt, and it +is the reason this section stays open rather than being struck out. + +**Switching is a reaction, not a declaration.** A licence that hits its limit and must be swapped is +a response to something observed. Expressing it as a declaration would make the declaration mean +*whatever is working right now*, which is not a thing anybody declared. It belongs with +observability, changing a binding — and the binding is then declared as usual. **Saying this +plainly is what stops the declaration language growing a conditional**, and nothing built here +grew one. + +## The vendor-agnostic generalisation + +*[ADR 0050](../../02-DECISIONS/0050-model-access-is-vendor-agnostic.md) is accepted; this section +describes what it decides. The section above stands as what is built today.* + +What runs is one vendor — the mesh's Anthropic feature. A read-only trace asked whether the +`model-access` provision is Anthropic-shaped or genuinely general, and found that the general layer +already exists: a licence is a record with a `vendor` and a non-secret `serves`, a holder's +credential is sealed per holder, `accept` takes an operator-supplied value and discards the +plaintext, and a locally-run model answers at node scope with no licence. None of that names +Anthropic. What is Anthropic's is a thin band: the credential is a subscription OAuth grant — an +hourly access token and a refresh token — and that shape alone drags central rotation, a +refresh-token-stripping delivery, an identity guard and a `utilization%` usage reading behind it. + +**A vendor is an adapter.** The vendor-specific lifecycle moves into a per-vendor adapter selected by +the licence's `vendor` field — the same shape as a registrar-scoped `public-dns` provider behind the +neutral `public-dns` interface ([ADR 0044](../../02-DECISIONS/0044-a-public-name-is-provisioned-like-any-capability.md)). +A consumer still requires `model-access` and never names a vendor; the interface is drawn at the +consumer's real coupling — *reach a model* — per [ADR 0040](../../02-DECISIONS/0040-what-a-module-is.md). +The field is `vendor` rather than `provider`, because "provider" already means *which node answers a +brokered provision* and the two facts must not share a word. + +The adapter declares a `shape` — `static-key` or `refreshable-grant` — and, optionally, the verbs a +vendor happens to need: `accept` a supplied key, `refresh` a grant, read an `identity` off the +credential to catch a mis-binding, report `usage`, and `deliver` the value. **A static-key vendor +implements almost nothing** — a supplied key, sealed to its holders, delivered. The abstraction is +built so the common vendor is small and the rare one carries its own weight. + +``` + requires: model-access (the consumer, vendor-blind) + │ + ┌──────┴───────┐ + │ a licence │ vendor: … serves: base URL, model + └──────┬───────┘ + selected by │ vendor + ┌─────────────────┼──────────────────────────┐ + ▼ ▼ ▼ + anthropic-api-key anthropic (another vendor) + shape: static-key shape: refreshable-grant + accept, deliver accept, refresh, identity, + usage, deliver +``` + +**The crux is one relaxation, stated plainly.** A refreshable credential cannot be both sealed so the +mesh cannot read it *and* rotated centrally — rotation needs a readable refresh token, and the working +central rotation is the half [ADR 0024](../../02-DECISIONS/0024-model-access-is-a-provision.md) keeps +on purpose. So for `refreshable-grant` vendors only, the **manager node holds the refresh token +encrypted at rest** — a bounded, declared exception. Access tokens stay sealed per holder, and the +refresh token is stripped on delivery, so *a node never holds a refresh token* remains true for every +node but the one manager. **Static-key vendors keep the full guarantee**: there is nothing to rotate, +so `accept` discards the plaintext and the carve-out never fires — and static-key is the majority. The +exception is written down because a relaxed guarantee that is not stated is indistinguishable from a +broken one, and it is narrow on three axes at once: refreshable-grant only, the refresh token only, +the manager node only. + +Usage is normalised to `(licence, consumer, period, metric, value)` with the raw response kept beside +it; the metric is vendor-defined, so no false common unit is forced. Anthropic becomes the first +`refreshable-grant` adapter, and `anthropic-api-key` — the same vendor's plain keys — is the +`static-key` case that proves the abstraction is more than one vendor in disguise. + +## How it is checked + +In the lab, on real machines, in the order a person would meet it: a consumer is refused with both +candidates named; put on one and still refused because no key exists; the key is given on standard +input and not echoed; the public half arrives saying it came from a record rather than a machine; +the key arrives readable only by that machine — and it is **nowhere in the control plane's own +database**, nor in anything that crossed the broker. + +For the generalisation ([ADR 0050](../../02-DECISIONS/0050-model-access-is-vendor-agnostic.md)): a +second vendor — `anthropic-api-key`, static-key — is bound to a consumer and exercises the whole path +with the carve-out switched off, its key sealed per node and absent from the control plane's database. +For a `refreshable-grant` licence, the refresh token is asserted to exist (encrypted) **only on the +manager node**, to be **absent from every holder's delivery**, and the delivered credential to be +access-token-only; a `static-key` licence stores no refresh token anywhere. A scenario with two +vendors confirms each licence's lifecycle runs its own adapter, selected by `vendor`. diff --git a/03-DESIGN/01-to-be/15-the-agent-session.md b/03-DESIGN/01-to-be/15-the-agent-session.md new file mode 100644 index 0000000..db1528b --- /dev/null +++ b/03-DESIGN/01-to-be/15-the-agent-session.md @@ -0,0 +1,226 @@ +--- +layer: to-be +status: designed +code: [mesh-control, mesh-host] +updated: 2026-08-31 +decisions: + - 02-DECISIONS/0004-a-node-and-how-it-joins.md + - 02-DECISIONS/0026-the-mesh-has-a-session-of-its-own.md + - 02-DECISIONS/0024-model-access-is-a-provision.md + - 02-DECISIONS/0025-the-design-record-is-read-not-copied.md +--- + +# The agent session + +**One mechanism, started twice.** A node's session and the mesh's session are the same thing +pointed at different context. This document describes the mechanism; where the two differ it says +so, and the differences are few enough to list here: + +| | node session | mesh session | +|---|---|---| +| **address** | the node's name | the mesh | +| **context root** | the node's | the mesh's | +| **engram** | that node's | the mesh's | +| **licence** | bound in its own right | bound in its own right | +| **runs on** | that node | the node holding the control plane | +| **how many** | one per node | one | + +Everything below applies to both unless it says otherwise. + +## What a session is, restated for what is being built + +[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) settled the behaviour: permanent, +remembering across callers, its own tools, reachable over the broker, and — switched off — still +answering *I am switched off* rather than falling silent. + +**Nothing said how one is set up**, and that gap is what this document closes. It was invisible +until [ADR 0026](../../02-DECISIONS/0026-the-mesh-has-a-session-of-its-own.md) needed to describe +a second instance, because *the same as that one, elsewhere* means nothing until the first has +been written down. + +## The context root is the whole of the difference + +**A session is defined by the directory it starts in.** That directory holds the engram, the +session's tools, and whatever standing instruction it works under. Two sessions differing only in +their root are two different agents, and nothing else has to differ to make them so. + +This is deliberately a **small** definition. The alternative — a session type, with the mesh +session as a distinct kind — would mean two implementations of one mechanism, and +[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) records what that +costs: two things nearly the same, built twice, until neither word means one thing. + +**The root is delivered as declared state, not carried by the session.** It is files on a machine, +which is precisely what the host applies +([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). A session's context therefore changes the +way anything else changes — the mesh declares it, the host writes it — and there is no second +mechanism for shipping an engram. + +**Changing the engram is changing a file.** So it is versioned, reviewable, and rolled back like +any other declared state; and a node whose engram was changed reports having applied it, the same +as it reports anything else. + +## Where each one runs + +**A node's session runs on that node**, and cannot be moved. Moved, one machine is answering as +another (ADR 0004). + +**The mesh's session runs on the node holding the control plane.** The reasoning is in ADR 0026 +and is worth carrying here because it is easy to get backwards: this is not *the important agent +goes on the important machine*. It is that the control-plane node is already the one place +excepted from *compromise of a node is compromise of that node*, and an agent able to reach +everything, placed anywhere else, would create a second such place. + +## Messages + +**A session is reached over the broker** +([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)). There is no second +transport, nothing is dialled at a session, and the mesh session is not a hop: node-to-node +messages continue to travel directly, and nothing is routed through it. + +**A message becomes a prompt; a reply travels back the same way.** Whoever asked — a person at a +surface, another session, a worker — is a caller, and the session remembers what each of them +asked, together, over time. + +**Being asked something it does not have, a session may ask another.** How it does so is its own +business and follows from its engram rather than from a message format: it may say who wants to +know, or simply ask. A person relaying a question makes the same choice. + +**A switched-off session answers.** The queue is still read and the state is the reply, with no +model involved. This is the same rule the host follows about a service that does not exist, and +it is the rule this repository has now paid for three times: **absence must never be +indistinguishable from a failure to answer** +([005](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md), +[008](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md)). + +## Model access is bound to the session, not to the machine + +**A session is a consumer in its own right.** This is the gap +[`14-model-access.md`](14-model-access.md) records as *a consumer that is not a machine*, and it +is the one part of this design that cannot be built from what exists: the provisions model has no +consumer identity other than a node, so today a binding can only say *this module on this +machine*. + +That is not sufficient here, and the shortfall is concrete rather than theoretical: + +- the control-plane node hosts **two** sessions, which must be able to hold **different** + licences — a per-machine binding cannot express it at all; +- *this node's session uses the personal licence, the mesh's uses the company one* is the ordinary + case, not an exotic one. + +**What is delivered still lands on a machine. What is chosen is chosen per session.** Delivery +follows the session's context root, which is where its credentials belong — the same rule ADR 0001 +states for an agent's config directory, with the root standing in for it. + +**Switching a licence remains a reaction, not a declaration** (ADR 0024). A session that exhausts +a licence is observed and its binding changed; the binding is then declared as usual. Nothing here +grows a conditional in the declaration language. + +## What it remembers, and where + +**A session's memory lives in its context root**, beside its engram and its tools. That is the +same rule as everything else here rather than a new one: the root is the whole of what makes one +session a different agent from another, and memory is part of what makes it *that* agent. + +**The mesh session's memory is its own.** It is not assembled from the node sessions on demand, +and it is not a view over theirs. What the mesh has been asked, and what it worked out, is held +in the mesh's root — not in the root of the node that happens to host it. + +**That distinction is the point of putting it there.** The control-plane node runs two sessions +on one machine. If memory belonged to the machine rather than to the root, they would share it, +and the mesh's recollection of a fortnight of questions would be indistinguishable from that +node's own — which is the collision ADR 0026 exists to avoid, arriving through the back door. + +**The root holds two kinds of thing, and confusing them destroys the memory.** The engram and the +tools are **declared**: the mesh says what they are and the host writes them, so editing one on +the machine survives until the next heartbeat and no longer +([ADR 0011](../../02-DECISIONS/0011-managed-files-are-generated-never-edited.md)). The memory is +**written by the session itself** and is declared by nobody — the mesh does not get to say what a +session remembers, and a mechanism that regenerates the root wholesale would erase a fortnight of +it on the next pass, silently, while reporting success. + +So the root is not uniformly managed, and **which parts are must be explicit rather than +inferred**. A session's memory is its own output, kept across restarts, backed up as data rather +than reproduced from a declaration — because there is nothing to reproduce it from. + +## What the mesh session knows + +**It holds the design record by reading it** +([ADR 0025](../../02-DECISIONS/0025-the-design-record-is-read-not-copied.md)), not by holding a +copy. It is that reader; there is not a second agent for it. + +**It answers into a symptom search**, so what it knows appears beside ordinary results rather than +only when it is asked. **And when it cannot be reached, the search says so.** A result set that +silently omits this material looks identical to one where nothing matched — the same rule as the +switched-off session above, at a different layer. + +**Reading is one-way.** It reads the repository and answers from it; nothing flows back. The +repository is public and the mesh is not, and a return path is how installation-specific detail +arrives into documents that must not carry it. + +## What it does with work + +**It dispatches; it does not hold a queue.** Asked for something that belongs elsewhere, it asks a +node session or hands the work to a worker. It is not an employee +([ADR 0003](../../02-DECISIONS/0003-agents-are-persistent-employees.md)): nobody hires it, it is +never drained, and it is not reassigned. + +The distinction is the one ADR 0001 was written to keep: a session comes with the thing it belongs +to and goes when that thing goes; a worker is hired, holds tasks, and moves. Built from the same +parts, run on entirely different terms. + +## The surface is separate + +That a person usually reaches the mesh session through a board is likely and is not settled here. +The session is reachable over the broker like everything else; what puts a text box in front of it +is a different design, and the session does not know which surface asked. + +**Several sessions open at once is a property of the surface, not of the sessions.** A board +showing the mesh's session beside one per node, switchable, leaves *one per node, permanent* +untouched: each is still one conversation remembering all its callers together. Only *concurrent +conversations with the same session* would touch ADR 0004, and that is not what is wanted. + +Such a surface is a **client** and holds nothing — it sends prompts and shows what the sessions +themselves remember, which is what keeps it inside `11-a-board.md`'s rule that the board is not a +second implementation of anything. Prior art to draw the interaction from: impire's *soulstream* +(already cited in [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md)) and +`herdrdev/herdr`. **Not yet designed, and not on the path to replacing provisioning** — recorded +here so the shape is not rediscovered. + +## Consequences + +**Two sessions on one node, and no ambiguity.** The control-plane node hosts its own node session +and the mesh session. ADR 0004's *one per node* forbids ambiguity about who answers when a **node** +is addressed; these answer to different addresses. + +**A new node gets a session by being declared, not by being set up.** The context root is declared +state, so a joining node's session arrives the way its packages and services do. + +**The mesh session is a single point of convenience, and must never become one of reach.** Every +node stays directly addressable. If that ever stops being true, the failure is quiet in the worst +way: every machine keeps working and nobody can ask about it. + +## How it is checked + +**A test names the decision it defends** ([ADR 0017](../../02-DECISIONS/0017-a-test-defends-a-decision.md)), +and these run in the lab on real machines: + +| Check | Defends | +|---|---| +| a node is asked something and its session answers | ADR 0004 | +| the mesh is asked something and the mesh session answers, on the control-plane node | ADR 0026 | +| both sessions on that node answer, to their own addresses, without ambiguity | ADR 0026 | +| a session switched off replies saying so, rather than timing out | ADR 0004 | +| a session whose engram was changed reports having applied it, like any declared file | ADR 0005 | +| the two sessions on one node hold different licences, and each uses its own | ADR 0024 | +| a node is messaged directly while the mesh session is stopped, and answers | ADR 0001 | +| a search consulting an unreachable mesh session says it was not consulted | ADR 0025 | + +The last two are the ones worth writing first. Both defend properties that are invisible while +everything works, and both describe a mesh that looks entirely healthy at the moment it has +stopped telling the truth. + +## Deliberately not decided + +**How a person's identity reaches a session.** Callers are distinguished, but who a caller *is*, +and whether a session should act differently for different people, is the human-agent question +ADR 0001 leaves open and this does not close. diff --git a/03-DESIGN/01-to-be/16-module-coverage.md b/03-DESIGN/01-to-be/16-module-coverage.md new file mode 100644 index 0000000..8b993fa --- /dev/null +++ b/03-DESIGN/01-to-be/16-module-coverage.md @@ -0,0 +1,176 @@ +--- +layer: to-be +status: designed +code: [] +updated: 2026-09-01 +decisions: + - 02-DECISIONS/0009-modules-and-the-graph.md + - 02-DECISIONS/0005-the-node-host.md + - 02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md +--- + +# What a module must be able to say + +**Measured, not guessed.** 127 manifests in the system being replaced were read and every key +counted, then set against what the new manifest can express. This document is the coverage +checklist: what is already sayable, what is deliberately not, and what is missing. + +*Surveyed 2026-09-01. Counts are modules, not occurrences, unless stated.* + +## Already sayable + +| what it says | used by | how it is said here | +|---|---|---| +| **depends on another module** | 65 | `requires` naming a module. A requirement naming a module means *that module*, not anything providing the name | +| **system packages** | 28 | the `package` shape | +| **a container** | 48 | the `container` shape, pinned by digest | +| **systemd units** | 17 | a `file` for the unit, a `service` for the state it should be in | +| **how to reach it** | 17 | `serves`, with the mesh adding which machine and where | +| **a public name** | 11 | requiring `route` and contributing the name | +| **ports it opens** | 11 | `listens`, from which filtering is computed | +| **data directories, and who owns them** | 38 | `directory` with `owner`; never removed while holding anything ([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)) | +| **what it provides and requires** | 15 + 2 | `provides` / `requires`, named for what the consumer is coupled to | +| **restart when something changes** | 11 | `restart-on` | +| **a generated credential** | 20 | `own-secrets`, sealed to the machine | +| **a value that differs per node** | 43 | settings, and `computed` for what only the mesh knows | +| **images built from source** | 2 | `build.artifacts` | + +## Deliberately not sayable + +**Stage hooks — 36 modules.** Arbitrary code at install, configure and start. **The link may not +carry an action** ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)): what may be pushed is +bounded by form, and a command to run is not a form. A module needing setup logic ships a program +that reads what the mesh delivered and reconciles — which is what the provisioners are, and they +are ~350 lines each including the reasoning. + +**Flavours — 6 modules.** Variants of one module. Retired in favour of claims +([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)): two display servers are two +modules that both claim the seat, and adding a third changes nothing anywhere else. What is lost +is `extends` chains, which were doing inheritance and are better as separate modules. + +## Missing, and what each would take + +Ordered by how many modules need it. The largest entry turned out not to be a gap at all, which +is left in place rather than deleted: the first framing of it was wrong in an instructive way, and +a checklist that quietly loses its biggest item reads as though nobody looked. + +### Tool servers — 56 modules — *not a gap* + +Over half the modules ship a `tools/` directory that becomes tools an agent can call on that node. +This was written up as the largest single gap. **It is expressible with what exists**, and the +first framing of it was wrong in a way worth keeping: *a module provides `tools`, the session +requires them* does not work, because a requirement has exactly one answer and 56 modules offering +tools would be 56 answers to one question. + +Turn it around and it fits exactly. The session **provides** `tool-host`; every module offering +tools **requires** it and **contributes** where its tools are. Many-to-one is what `contributes` +has always been, and the session receives all of them in one file: + +``` +given: { from: gitea, values: { at: … } } + { from: minio, values: { at: … } } + { from: umami, values: { at: … } } +``` + +Verified by resolving it, not by reading the code. It also only became possible today: until +[`022`](../../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md) was +fixed, several modules on one node requiring the same thing was refused outright. + +What remains is not vocabulary but a decision about **what a tool server is** — a container the +module already runs, and what the session does with the list. That is work, not a missing shape. + +### Schema migrations — 14 modules + +A module with a database needs its schema brought up to date before it runs. The mesh does this +for its own contexts and has no way for a *module* to declare it. The provisioner pattern covers +it — a program that runs migrations and exits — but nothing expresses *this must happen before +that starts*, which is the actual requirement. + +### Configuration merging — 18 modules, 134 files + +Files assembled from a module's default plus per-node overrides, with a strategy (`replace`, +`merge`) and a format (`toml`, `yaml`, `json`). Settings already merge into a file's content; what +is missing is format-aware merging. + +**And it should stay missing.** A mechanism that understands TOML will be asked for YAML, then +INI — which is how the arrangement being replaced became something nobody could hold in their +head. The module knows its own format because it wrote the rest of the file. + +### Health checks — 7 modules + +`{type: port|url, expect: …}`. The mesh knows whether a container is running, which is not the +same as whether it answers — a distinction this project has paid for twice already. + +An `action` carries a `verify` and is exactly this shape. **It is not available to a module**: the +link may not carry a command to run ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)), and a +module's resources reach a machine over the link. So a health check needs a way to say *ask this +and expect that* without saying *run this* — closer to a `listens` entry than to an action. + +**This is the gap most worth closing**, because *running* and *answering* being conflated is a +class of fault, not an inconvenience. + +### Theme knobs — 3 modules, 101 values + +`{theme: {kind: color|font|string, label}}` — declared so a ricing tool can offer them. Settings +already carry the value; what is missing is the **metadata** saying a value is presentable and +what kind it is. Small, self-contained, and only interesting once something presents them. + +### Event routing — 2 modules + +`{routing-key: tool}`, generating a consumer. Two modules; wait for a third before deciding. + +### Publishing a package — 6 modules + +Modules published to a registry and consumed as libraries. This is a *build* output the mesh does +not deliver to a node, so it may not belong here at all. + +## What checking the coverage found + +Two faults, both surfaced by asking what a *real* node looks like rather than what a test does. +Neither is about the vocabulary; both are about the machinery under it. + +**A credential belonged to a machine, not to a module** +([`022`](../../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md), +fixed). A node running three services against one database could not be planned at all — and on +the consumer's side did not refuse, it just gave two of the three no credential. Every scenario +written to date had one consumer per node, which is the natural shape of a small test and not the +shape of a machine. + +**A consumer could not build a connection string** +([`023`](../../04-ISSUES/023-a-consumer-cannot-build-a-connection-string/00-report.md), fixed). It +had its password in the right shape and the host, port and user name were out of reach: the user +name was invented by the provisioner and recorded nowhere, and the bound values sat in a JSON +document that an application reading `KEY=value` cannot use. + +The asymmetry was backwards, which is what made it worth stating. **The secret is the hard case** — +the mesh must not be able to read it — and the secret was the part that already arrived. The host +and port are ordinary facts held in the clear, and they were the ones stuck. Both halves came from +the same thing: the mesh knew something and did not say it. + +## What the survey found that is not about coverage + +**Declaration and reality had drifted in the system being replaced.** Several live provisions are +brokered by modules whose manifests declare nothing — a speech-to-text engine served to a consumer +on another node, an object-store bucket held by a module whose manifest mentions only its +database. **A manifest that does not have to be true stops being true**, which is the argument for +resolution refusing rather than warning. + +**Two derivations of the same fact.** A module's kind was computed in two places from different +evidence — one from the manifest, one from what is on disk — producing different labels for the +same module. There is one derivation here, and there should stay one. + +**A live listing returned credentials in plaintext.** Not a coverage question, but the reason +sealing is worth its inconvenience. + +**An image store is a module, and was written up here as something the mesh does.** It was +considered for the substrate and removed, because the test is not *can it grant itself one* — +nearly anything passes that — but whether the control plane needs it before it can give its first +instruction. It does not. So a registry somebody runs for their own images is the same module as +the one the mesh runs for its own: it offers a place to push, and claims that role once per +machine. + +**A rule was enforced only at the far end.** A module may not declare an action, and the host +refused one correctly — but the control plane accepted it into the catalogue, resolved it and +pushed it, so the refusal arrived on a machine with nothing tying it back to the manifest. The +rule held; it was just unusable, which is the same shape as the network shape that cost five +failing tests before anyone read the host's log. It is now refused where it is written. diff --git a/03-DESIGN/01-to-be/README.md b/03-DESIGN/01-to-be/README.md index c9255d2..fda651a 100644 --- a/03-DESIGN/01-to-be/README.md +++ b/03-DESIGN/01-to-be/README.md @@ -9,18 +9,33 @@ document is written and this one's status becomes `implemented`. | Document | Covers | Rests on | |---|---|---| -| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) | -| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) | -| [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) | -| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) | -| [`04-lab-installation.md`](04-lab-installation.md) | Getting the lab onto a clean machine, and why it verifies capability rather than installation | [ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) | -| [`05-the-node-host.md`](05-the-node-host.md) | Tier 0 — the one thing installed by hand, and the only thing that changes a machine | [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) | +| [`00-work-breakdown.md`](00-work-breakdown.md) | How modules move across one at a time, until the old registry can be switched off | [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md), [ADR 0016](../../02-DECISIONS/0016-the-lab.md) | +| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-the-lab.md), [0029](../../02-DECISIONS/0016-the-lab.md) | +| [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0016](../../02-DECISIONS/0016-the-lab.md) | +| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0016](../../02-DECISIONS/0016-the-lab.md) | +| [`04-lab-installation.md`](04-lab-installation.md) | Getting the lab onto a clean machine, and why it verifies capability rather than installation | [ADR 0010](../../02-DECISIONS/0010-delivery.md) | +| [`05-the-node-host.md`](05-the-node-host.md) | Tier 0 — the one thing installed by hand, and the only thing that changes a machine | [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) | +| [`06-the-control-plane.md`](06-the-control-plane.md) | Tier 2 — what the term means, and the test for what belongs in it | [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) | +| [`07-the-substrate.md`](07-the-substrate.md) | Tier 1 — what the control plane consumes and cannot grant itself | [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md), [0048](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) | +| [`08-connectivity.md`](08-connectivity.md) | One context in full — overlay, resolution, exposure, filtering, certificates | [ADR 0007](../../02-DECISIONS/0007-connectivity.md), [0050](../../02-DECISIONS/0007-connectivity.md), [0051](../../02-DECISIONS/0004-a-node-and-how-it-joins.md), [0055](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) | +| [`09-the-node-lifecycle.md`](09-the-node-lifecycle.md) | How a machine becomes a node, stays one, and stops being one | [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md), [0051](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) | +| [`10-delivery.md`](10-delivery.md) | Modules, the three edges, and how a change becomes a running thing | [ADR 0010](../../02-DECISIONS/0010-delivery.md), [0064](../../02-DECISIONS/0009-modules-and-the-graph.md), [0065](../../02-DECISIONS/0009-modules-and-the-graph.md) | +| [`11-a-board.md`](11-a-board.md) | What a person sees of the mesh, and why it is read from what runs | [ADR 0008](../../02-DECISIONS/0008-a-context-owns-its-store.md), [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) | +| [`12-a-module-repository.md`](12-a-module-repository.md) | A module repository, and what builds it | [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md), [ADR 0010](../../02-DECISIONS/0010-delivery.md), [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) | +| [`13-credentials-and-their-rotation.md`](13-credentials-and-their-rotation.md) | Credentials, and moving them without a consumer holding one the provider does not know about | [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md), [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) | +| [`14-model-access.md`](14-model-access.md) | Model access as a provision, and what a licence is bound to | [ADR 0024](../../02-DECISIONS/0024-model-access-is-a-provision.md), [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) | +| [`15-the-agent-session.md`](15-the-agent-session.md) | One mechanism started twice — a node's session and the mesh's | [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md), [ADR 0026](../../02-DECISIONS/0026-the-mesh-has-a-session-of-its-own.md) | +| [`16-module-coverage.md`](16-module-coverage.md) | What a module must be able to say, measured against 127 that exist | [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md), [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) | ## Not yet written -- **The eight bounded contexts.** [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) - decides the decomposition; the per-context specifications do not exist yet. The work - breakdown says in what order they are needed. -- **Domain grouping outside the core.** [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) - settles the principle and explicitly does not settle the domain list. That is a research - effort, not a design document, until it concludes. +- **The remaining six contexts.** + [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) + settles the list at seven; `connectivity` is the first written in full + ([`08`](08-connectivity.md)) and the other six do not exist yet. The work breakdown says in + what order they are needed. +- ~~**Domain grouping outside the core.**~~ **Not needed.** + [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) is + superseded by [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md): + there is no domain module to group into, so there is no domain list to settle. Relationships + are edges, and grouping is a tag and a query. diff --git a/04-ISSUES/001-failed-package-install-reports-success/00-report.md b/04-ISSUES/001-failed-package-install-reports-success/00-report.md index 02a0655..0b61223 100644 --- a/04-ISSUES/001-failed-package-install-reports-success/00-report.md +++ b/04-ISSUES/001-failed-package-install-reports-success/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-08-22 -located-in: [] -fixed-by: +located-in: [mesh-host] +fixed-by: mesh-host — a package is read back from the package database after installing amended-design: --- @@ -31,14 +31,14 @@ exists to catch — a step that failed, reported success, and left the next step state that was never produced. It is also a direct violation of a decision already taken and recorded: -[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) says a step that fails must fail the +[ADR 0010](../../02-DECISIONS/0010-delivery.md) says a step that fails must fail the job. That record notes the rule is applied instance by instance and enforced by no mechanism. This is an instance where it was never applied. ## Evidence - Observed 2026-08-22 while declaring the virtualisation package required by - [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md). + [ADR 0016](../../02-DECISIONS/0016-the-lab.md). - A fix is written and open as a pull request, unmerged since 2026-08-20. ## Open questions @@ -47,3 +47,18 @@ This is an instance where it was never applied. - Is this specific to package installation, or does the surrounding stage swallow every non-zero exit? - The fix has been open for two days. What is the review path for a change of this class? + +## How it is answered + +*2026-08-31.* **The host reads the package database back after installing**, and refuses when it +does not have the package: + +> ` was installed without error and the package database does not have it` + +That is the general rule this issue is one instance of, and the host applies it to everything it +does: a command exiting zero says a transaction was *accepted*, not that the machine changed. The +same read-back is why a container that starts and immediately dies fails an apply, and why a +service asked to run is checked rather than assumed. + +HAL keeps the fault until its provisioning is switched off. Fixing it there would mean +implementing the read-back twice, in the system being replaced. diff --git a/04-ISSUES/002-stale-package-index-fails-silently/00-report.md b/04-ISSUES/002-stale-package-index-fails-silently/00-report.md index ccbc7b7..a9bfb5e 100644 --- a/04-ISSUES/002-stale-package-index-fails-silently/00-report.md +++ b/04-ISSUES/002-stale-package-index-fails-silently/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-08-22 -located-in: [] -fixed-by: +located-in: [mesh-host] +fixed-by: mesh-host — a stale package index is named rather than reported as a failed install amended-design: --- @@ -38,3 +38,29 @@ the job is green. versions — or is an index sync part of the install step? - A partial sync is unsafe on the platform in use; a full upgrade is the only sanctioned fix. Does that make index freshness a scheduled node concern rather than a pipeline one? + +## How it is answered + +*2026-08-31. It was present in the replacement too, which is why this is a fix rather than a note +saying the new mesh does not have it.* + +**The failure is named.** A machine asking for a version the mirrors have replaced now says so, and +says what fixes it — a full upgrade of the machine. + +**It is deliberately not fixed by synchronising.** `pacman -Sy ` installs a package built +against libraries the machine does not have: a partial upgrade, unsupported on this distribution, +which surfaces much later as something apparently unrelated. That is a decision about the whole +machine, and a host that made it silently while applying one resource would be taking a large +decision in a small place. + +So the host distinguishes the two cases and leaves the decision where it belongs. **A declaration +that is wrong and a machine that is out of date fail identically otherwise, and they are fixed in +completely different places.** + +**And the package manager's own words were being thrown away** — the output was read into `_`, so +the 404s that name the cause never reached anybody. Whatever it said is now part of the failure, +which is the rule everywhere else here and was not being followed in the one place where the reason +exists only in the output. + +*Checked by a stale-index failure being named as one, an ordinary missing package not being, a +single mirror timing out not being, and a successful install still saying nothing.* diff --git a/04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md b/04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md index f2b5d5c..7b670e1 100644 --- a/04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md +++ b/04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md @@ -1,9 +1,9 @@ --- -status: open +status: resolved opened: 2026-08-22 -located-in: [] -fixed-by: -amended-design: +located-in: [mesh-control] +fixed-by: mesh-control — a machine's filtering is computed from what it was assigned +amended-design: 03-DESIGN/01-to-be/08-connectivity.md --- # 003 — A firewall rule's `scope:` is read by no code @@ -38,3 +38,29 @@ any check. - Were the five declarations intended to restrict something that is currently open? Each needs checking against what the node actually exposes — the declaration cannot be trusted either way. + +## How it is answered + +*2026-08-31.* Both halves, in Novox Mesh. HAL keeps the fault until its provisioning is switched +off, which is what this issue is now waiting on rather than a fix of its own — patching `scope:` +into something that works would mean implementing it twice, in the system being replaced. + +**The unknown key.** A manifest is parsed strictly: an unknown key is refused with the key named, +the discipline the host's declaration parser has always had. `scope:` would not survive being +written today, and neither would a misspelling of anything else. This is the general fix — the +issue's own observation was that *any* invented key behaved this way, and that the one instance +was found by reading rather than by any check. + +**The rule that restricts nothing.** `scope:` is not reimplemented. A module says what it listens +on and **who may reach it**, and saying from where is required rather than defaulted: a rule with +no source is open, and must say so rather than appear to restrict something. A machine's whole +rule set is then derived from every module assigned to it — so there is no second list to keep in +step, which is the condition that let the first one drift out of use unnoticed. + +**And it is enforced, which is the part that makes this different from before.** The mesh renders +the rule set; a service on the node is declared to reflect that file, so replacing it restarts +what loads it. Proven in the lab against two real ports on a real machine: the declared one +answers from another machine, the undeclared one does not, and removing the module that wanted the +port closes it with nobody editing a rule. + +The design is [`03-DESIGN/01-to-be/08-connectivity.md`](../../03-DESIGN/01-to-be/08-connectivity.md) §4. diff --git a/04-ISSUES/004-certificate-issuance-targets-production/00-report.md b/04-ISSUES/004-certificate-issuance-targets-production/00-report.md index 23f7c51..f1dc315 100644 --- a/04-ISSUES/004-certificate-issuance-targets-production/00-report.md +++ b/04-ISSUES/004-certificate-issuance-targets-production/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-08-22 -located-in: [] -fixed-by: +located-in: [hal] +fixed-by: hal — ACME_CA_SERVER selects the authority, and defaults to staging amended-design: --- @@ -21,7 +21,7 @@ recoverable by retrying — it removes the ability to issue a certificate anyone The consequence lands hardest on exactly the work most likely to iterate: standing up a new node, changing how names resolve, or testing the lab's certificate authority split -([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)). +([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). ## Evidence @@ -35,3 +35,41 @@ node, changing how names resolve, or testing the lab's certificate authority spl - The lab issues its own certificates and so does not consume public quota at all. Does that make this a problem only for experiments run outside the lab, and therefore an argument for running them inside it? + +## Resolution + +*2026-08-31.* **The authority is now selectable, and the default is staging.** + +Confirmed first: the resolver declared no `caServer` at all, so the client fell to its built-in +default of the public authority's production endpoint. There was no setting to change — not a +setting set wrongly. + +`ACME_CA_SERVER` now names the authority, and **defaults to staging**. That answers the first +open question in the affirmative: it is a node property, and the node that serves real traffic is +the one that states so. + +**Why staging is the default rather than the safe-looking alternative.** Defaulting to production +and documenting the override would leave the safe path depending on somebody remembering to opt +out of it — on exactly the work most likely to iterate. That is the same fault as +[005](../005-pipeline-test-harness-unbuildable/00-report.md): *remembering is not a mechanism.* +A staging certificate is trusted by no browser, so the mistake announces itself in the first +request rather than a fortnight later when the quota is gone. **The failure that is loud and +immediate is the cheaper one**, and quota exhaustion is neither. + +**The second open question is answered too, and it is not the whole answer.** The lab issues its +own certificates and consumes no public quota, so experiments belong there. But "run it in the +lab" is advice, and the nodes this issue is about are the ones outside it — the default is what +protects those. + +### The rollout is ordered, and the order is the dangerous part + +Both public-serving nodes were checked: neither set the variable. Applying the change without +pinning them first would re-issue their public certificates from an untrusted authority and break +TLS for every hosted service — **this issue's own failure mode, arriving through its fix.** Pin +first, then merge. The commit carries the exact commands. + +## Deliberately not done + +**Nothing was changed on a running node.** Pinning the public nodes and regenerating their +environment restarts the reverse proxy that fronts every hosted service, and that is an operator's +decision rather than a fix's. diff --git a/04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md b/04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md index 1957ffe..def605b 100644 --- a/04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md +++ b/04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md @@ -1,9 +1,9 @@ --- -status: open +status: resolved opened: 2026-08-22 -located-in: [hal] -fixed-by: -amended-design: +located-in: [hal, mesh-lab] +fixed-by: mesh-lab — a run leaves a receipt, and the receipt says what it covered +amended-design: 03-DESIGN/01-to-be/01-end-to-end-testing.md --- # 005 — The end-to-end pipeline harness has not built since the workspace was removed @@ -29,7 +29,7 @@ coverage was assumed, not checked. ## Evidence - The workspace was removed by pull request #240 on 2026-06-04 - ([ADR 0007](../../02-DECISIONS/0007-no-npm-workspace.md)). + ([ADR 0014](../../02-DECISIONS/0014-no-npm-workspace.md)). - The harness has not built since that date. - Recorded in the knowledge base as a standing entry, not as a fixed incident. @@ -45,3 +45,52 @@ unmentioned: until the lab exists, this is the coverage the pipeline is presumed - Repair, or retire in favour of the lab? Leaving it in the repository unbuilt is the one option that keeps the false impression of coverage. - Was anything relying on it, or had it already stopped running before the workspace removal? + +## Resolution + +*2026-08-31.* **Retired in favour of the lab, and the reason it went unnoticed was fixed +separately from the harness itself.** + +The old harness is not repaired. What replaced it is the end-to-end suite on a lab mesh, which +raises real machines and proves the pipeline against them. That answers the first open question. + +The second finding is the one worth keeping. *Nothing runs it, and nothing reports that nothing +runs it* is not a fact about that harness — it is a fact about **any** suite too expensive to run +on every push, and the lab suite is exactly that: it needs a machine with a hypervisor, so it runs +when somebody remembers. **Remembering is not a mechanism**, and the replacement inherited the +fault it was replacing. + +So three things now hold, each checked by a test that was confirmed to fail without it: + +- **A run leaves a receipt** — when it ran, what passed, and the commit each repository was at. + Kept outside version control, because the question is *has this machine run it*, and a receipt in + git would be a claim about everybody's machine made by whoever committed last. +- **The receipt can be judged, and says why it does not count.** Old, failed, taken against code + the repositories have since moved past, or a run that never raised a machine — each reads + differently, and only the last of those is new. **A receipt that says nothing about something is + not a receipt that clears it.** +- **The artifacts are rebuilt by the run, not beside it.** The suite consumes three artifacts from + two repositories. They were rebuilt by hand, from memory, and a rename that needed two of them + got one — leaving a binary eleven hours old refusing a field the mesh had just renamed, found by + a full run. That step now lives in the repository rather than in a terminal history. + +### What this issue taught twice + +**The fix reintroduced the fault, in miniature, and the second time was caught by running it.** + +The suite takes paths, so it can be pointed at one quick unit file — and the receipt from that run +was, at first, indistinguishable from a receipt for the real thing. A green record standing for a +run that raised no machines is this issue's own symptom, rebuilt inside its remedy. The receipt now +records what it ran, and a run that did not include the end-to-end file is not coverage. + +Separately, the code that decides *no receipt rather than a guessed one* — the rule that keeps the +record meaning something — was first written where no test could reach it. Writing "0 failed" +because nothing said otherwise is how a green record comes to mean nothing. + +And the counting itself **passed every test while reading nothing**: the test runner colours its +summary even into a pipe, so the anchored pattern never matched, and the fixtures it was checked +against were output that had been imagined rather than captured. **A fixture that agrees with the +mistake proves the mistake.** It is now checked against the runner's real bytes. + +Each of these was found by running the thing, not by reading it — which is the same argument this +issue makes about the pipeline. diff --git a/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md b/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md index a18403a..7c8280e 100644 --- a/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md +++ b/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md @@ -1,9 +1,9 @@ --- -status: open +status: located opened: 2026-08-23 located-in: [hal, hq] fixed-by: -amended-design: +amended-design: 02-DECISIONS/0025-the-design-record-is-read-not-copied.md --- # 006 — This repository is not indexed into the knowledge base, and the claim that it is holds up a decision @@ -59,7 +59,7 @@ checked it — including in the same commit that wrote the rule. ## Proposed direction — Nox is the search *Added 2026-08-23.* Rather than syncing these documents into the knowledge base, **Nox -([ADR 0027](../../02-DECISIONS/0027-the-product-is-novox-mesh.md)) works from within this +([ADR 0019](../../02-DECISIONS/0019-how-this-repository-works.md)) works from within this repository and holds its knowledge directly.** Retrieval becomes an agent reading the source, not a copy living in a second store. @@ -88,3 +88,70 @@ consults Nox — the answer is yes and the original promise holds. That is a design question for Nox, not a defect in this repository, and it should be settled before ADR 0019 is treated as answered. + +## Where this stands + +*2026-08-31. Re-checked, and deliberately not closed.* + +**The indexing still does not exist.** Two searches today, against both the symptom-indexed +memory and the structured archive, using a decision record's full title and a distinctive phrase +from a design document: no results, no partial match, no stale copy. The symptom in this report +is unchanged. + +**But the part that made it an issue is gone.** This report's argument was that the claim was +*load-bearing* — that a decision rested on a mechanism nobody had checked. It no longer rests on +it. The README now names the gap in the place the claim used to sit, and says it is left standing +rather than quietly reworded. The decision record that separates this repository does not invoke +indexing at all; its reasoning is cadence, reviewers, and scope, none of which depend on it. + +So what remains is not a false claim. It is an unbuilt capability and an open design question, +and those are different things. + +### What was done + +**A signpost, in the knowledge base, pointing here** — what lives in this repository, which +folders hold what, and when to come looking rather than search there. Explicitly a pointer and +not a copy: a derived copy drifts, and the enforced copy wins while the reasoned one quietly +stops being true. + +**It was tested, and it half works.** A search for *design records, decisions, repository* returns +it. A search phrased the way somebody would actually ask — *why is the mesh built this way* — +returns nothing, because the store matches terms rather than meaning. + +That is this report's own distinction, confirmed by measurement rather than argued: **a signpost +is reachable, it is not surfacing.** Someone who suspects the answer exists will now find it. +Someone debugging an error, with no reason to think this repository knows anything about their +symptom, still will not. + +### Why it stays open + +The question this report narrows to is unchanged and unanswered: + +> When a symptom is searched and the answer happens to live in a design document or a decision +> record here, does the searcher find it without already suspecting it exists? + +Today: **no.** Closing this means choosing between a one-way sync into the knowledge base and an +agent that reads this repository and contributes to a symptom search — and that is a decision +about how the knowledge system works, not a defect to be fixed quietly. + +**Marking it resolved while the indexing does not exist would be the failure this repository was +created to name**, one folder away from where it names it. + +## The direction is decided + +*2026-08-31.* **The agent reads this repository; nothing is copied.** Recorded as +[ADR 0025](../../02-DECISIONS/0025-the-design-record-is-read-not-copied.md), which also amends +what [ADR 0019](../../02-DECISIONS/0019-how-this-repository-works.md) promised: these documents +will not be *indexed*, they will be *read*, and the search consults the agent so its answers +appear beside ordinary results. + +A sync was the option that works with what exists today, and it was rejected on the one ground +this repository can least afford: it makes a second copy, and *the copy that is searched quietly +stops matching the copy that is edited*. + +**So the open question above is answered, and this report stays open on the build.** What closes +it is the check ADR 0025 names — search the mesh's memory for a phrase that appears only in a +design document here, and get it back. That check fails today by design. + +**What stands until then** is the signpost, and the honest description of it: reachable, not +surfacing. diff --git a/04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md b/04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md index 646bb1d..011eb1d 100644 --- a/04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md +++ b/04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md @@ -44,7 +44,7 @@ The distance between the two is the same one the delivery layer already has a na ## Why it matters now This is the first requirement of the lab -([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)), which is +([ADR 0016](../../02-DECISIONS/0016-the-lab.md)), which is phase 0 of the entire migration. The first capability the new work depends on is present, declared, and unusable — and would have stayed unusable silently. diff --git a/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md b/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md new file mode 100644 index 0000000..8839275 --- /dev/null +++ b/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md @@ -0,0 +1,102 @@ +--- +status: resolved +opened: 2026-08-28 +located-in: [hal] +fixed-by: hal — the script says it is manual, because it is +amended-design: +--- + +# 008 — The documented automatic node rescue does not exist + +## Symptom + +The mesh's documentation describes an automatic node rescue: a node that fails is recovered +without anybody intervening. **Nothing implements it.** + +Found incidentally while investigating supervision +([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what +actually supervises what: + +- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up; +- **nothing calls the rescue script on a timer**, so it runs only when a person runs it. + +The script exists. The thing that would invoke it does not. + +## Why this is worse than having no rescue + +A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a +gap nobody looks for, because the documentation says it is covered — and it is read exactly when +a node has failed and somebody is deciding whether to intervene. + +This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable +from a wrong one, and costs more, because people believe it.* Here the belief is that a failed +node recovers itself. + +## Scope + +**The as-is only.** The design being built has a different answer: +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts recovery +in a launcher that supervises the host, and that recovery is tested — 32 assertions, each +confirmed to fail when the behaviour is removed. + +So this issue is about the mesh that runs **now**, and it has two possible resolutions rather +than one: + +1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host + reaches the fleet. +2. **Delete the documentation** — and say plainly that a failed node needs a person, which is + what is true today. + +**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host +is, which is a scheduling question rather than a technical one. + +## What it would take to be sure + +Read back rather than assumed +([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)): list every unit on a +node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above +came from reading the repository, and confirming it against a running node is the difference +between *no unit declares this* and *no unit in the source declares this*. + +## Resolution + +*2026-08-31.* **Resolved the second way: the documentation now says what is true.** + +### Read back from running nodes, and the finding sharpened + +This report was written from the repository. Checked against three running nodes, as the section +above asks — and one of its own claims was wrong in a way that matters: + +| Claim | Verified | +|---|---| +| no unit declares `OnFailure=` | **true**, and in *both* scopes — of 604 system units, the only two are stock (`snapd`, `local-fs.target`); across ~30 user units, none | +| nothing calls the rescue script on a timer | **the timer exists.** `hal-health.timer` fires every 30 minutes | +| the rescue is not actually triggered | **true** — the *deployed* `hal-health.sh`, not only the repository's copy, contains no call to it | + +**The trigger exists and does not do the thing the script says it does.** That is worse than the +absence this report described, because it survives a halfway check: somebody verifying "is there a +health timer?" finds one, and stops. + +The precise falsehood was a single line in the rescue script — *triggered automatically by +`hal-health.timer` when the mesh service is failed.* It is now the opposite sentence, and says +what does trigger it. + +**Two further claims were found and narrowed.** Documentation in two places called the mesh +*self-healing* — once for peers refreshing a tool cache, once for nodes auto-updating on a repo +change. Both behaviours are real and neither is healing. **A phrase that overstates by a category +is read in the crisis it describes**, and "the mesh is self-healing" is exactly the sentence that +stops somebody intervening. + +### Why not the first way + +Implementing it was the other honest option, and it was not taken. The replacement host already +supervises recovery and that recovery is tested; adding unattended download-and-restart behaviour +to the fleet it is replacing is a change with real failure modes of its own — a rescue loop that +thrashes is worse than a node that waits. + +**This is the scheduling judgement this report said the choice turned on, and it is recorded +rather than assumed.** If rescue is wanted on the current fleet before the new host reaches it, +that is a decision to take deliberately: an `OnFailure=` and a call from the health script, both +small, neither free. + +Until then the documentation is true, which is the part that was costing something. diff --git a/04-ISSUES/009-a-digest-pinned-image-cannot-be-placed-in-the-lab/00-report.md b/04-ISSUES/009-a-digest-pinned-image-cannot-be-placed-in-the-lab/00-report.md new file mode 100644 index 0000000..b90a554 --- /dev/null +++ b/04-ISSUES/009-a-digest-pinned-image-cannot-be-placed-in-the-lab/00-report.md @@ -0,0 +1,114 @@ +--- +status: fixed +opened: 2026-08-28 +located-in: [mesh-lab, mesh-host] +fixed-by: + - "mesh-lab: a registry raised inside the scenario. Verified in a sealed machine — all four shapes applied with the image pinned by digest, idempotent, read back from the machine." +amended-design: +--- + +# 009 — A digest-pinned image cannot be placed in the lab, so `container` cannot be tested there + +## Symptom + +Two accepted decisions collide, and the collision makes one resource shape untestable. + +- **[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)** pins images by + digest, and the host **refuses** an image reference that is not pinned: + + ``` + resource "store": image "alpine:3.20" is not pinned. Write it as name@sha256:... + ``` + +- **The lab cannot place a digest-pinned image.** A sealed scenario cannot reach a registry, so + the lab exports an image from the workstation and loads it in the machine — and that loses the + digest. + +So a `container` resource is refused by the host if it names a tag, and unusable if it names a +digest. **There is no declaration the lab can currently raise that exercises the shape.** + +## What was measured + +Not inferred. `docker save alpine@sha256:d9e8…` produces an archive with **no repo tag**, because +a repo digest exists only for an image a registry served. Loading it says: + +``` +Loaded image ID: sha256:63f227… (not "Loaded image: alpine:3.20") +``` + +and `docker images` then lists nothing — the image is there but dangling. A container declaring +that digest therefore falls through to the registry: + +``` +Unable to find image 'alpine@sha256:d9e8…' locally +dial tcp: lookup registry-1.docker.io: no such host +``` + +which is correct behaviour on a machine with no route out. + +## What is not affected + +Everything else placed in the same sealed machine works, and was verified there: + +| shape | | +|---|---| +| `package` | applied, idempotent | +| `service` incl. `boot: enabled` | applied, read back as `enabled` | +| `action` | ran, verified | +| `container` | **blocked by this issue** | + +## Why it matters more than one shape + +The container shape is the substrate. Every step of raising a mesh past the container runtime is +a container ([`07-the-substrate.md`](../../03-DESIGN/01-to-be/07-the-substrate.md)), so the +bootstrap cannot be tested end-to-end until this is resolved — which is the thing the lab exists +for. + +## How it was fixed + +*2026-08-29.* A registry inside the scenario, as below — and it turned out to be the shape the +resolution predicted rather than a compromise on it. + +A scenario declares `images:` by tag. The lab stocks a registry **on the workstation**, where +there is a network, then raises one **inside the scenario** as scenery and serves them from it. +What a declaration pins is reported when the scenario is raised, because the digest belongs to +that registry and is not knowable before it exists. + +**The digests are the lab registry's own, and that is correct rather than a workaround.** What +[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) requires is a +reference that is exact and cannot move. A digest this registry assigned is both. + +Verified in a machine confirmed to have no route out: `package`, `service` including boot state, +a `container` pinned by digest, and an `action` inside that container — applied, idempotent on +re-apply, and read back from the machine rather than from the apply's own report. + +**One fault is worth keeping**, because it is this repository's own subject arriving in the +tooling built to catch it. The read-back checked that the registry's catalog endpoint answered, +by looking for the substring `repositories` — which `{"repositories":[]}` also contains. So it +**passed on a registry holding nothing**, and the failure surfaced much later as a container that +could not be pulled, a long way from its cause. It now asks for each image's manifest **by +digest**, which is what a machine actually does. + +## The shape of a resolution + +**A registry inside the scenario**, on its public segment, that machines pull from. That is not a +workaround: it is what the real mesh does — [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) +names an OCI registry as substrate, and every node after the first pulls from the mesh's own. +Testing against a registry is testing the real path rather than a stand-in for it. + +It also removes the lab's export-and-push mechanism rather than fixing it, which is the better +outcome: pushing image tarballs over the hypervisor was always a lab-only invention. + +**Not decided here**, because it is design rather than repair: where the registry runs, whether +it is scenery like the router ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) +or a placed artifact, and how images get into it. + +## Incidental, and already fixed + +The lab's own check on the load was too weak: it matched `"Loaded image"`, which is a prefix of +both `Loaded image:` and `Loaded image ID:`. So a load that produced an unusable dangling image +**reported success**, and the failure surfaced later as a container that would not start. It now +matches `Loaded image:` exactly and says what the runtime actually said. + +That is this repository's own subject arriving in its own tooling: a check that passes on the +wrong thing is worse than no check, because it moves the failure away from its cause. diff --git a/04-ISSUES/010-the-first-declaration-destroys-the-substrate/00-report.md b/04-ISSUES/010-the-first-declaration-destroys-the-substrate/00-report.md new file mode 100644 index 0000000..a477582 --- /dev/null +++ b/04-ISSUES/010-the-first-declaration-destroys-the-substrate/00-report.md @@ -0,0 +1,104 @@ +--- +status: fixed +opened: 2026-08-29 +located-in: [mesh-host, mesh-control] +fixed-by: + - "mesh-host: the store records where each resource came from — carried or declared — and each origin removes only its own. Verified in the lab on the exact scenario that caused this: the substrate survived, and a later declaration still removed what it had itself declared." +amended-design: +--- + +# 010 — The first declaration a node receives destroys the substrate it raised + +## Symptom + +A first node was raised from its carried bundle: container runtime, store, two context databases, +their schemas, broker, control plane. Eleven resources, all running. It then enrolled against the +control plane on its own machine, held its link open, and was sent a declaration naming two +resources — a directory and a file. + +Both were applied correctly. And **every container on the machine was removed**: the store, the +broker, and the control plane that had sent the declaration. The link died mid-sentence with +`the link closed: Exception (501) Reason: "EOF"`, because the broker carrying it had just been +torn down by the message it carried. + +Afterwards `mesh-host owned` listed two resources. The mesh had deleted itself. + +## What is actually wrong + +Nothing in the code is behaving incorrectly. `apply` removes what the store holds and the incoming +declaration does not name, which is what reconciliation means — the declaration is the desired +state, not a patch, and anything else would make it impossible to remove a resource by omission. + +**The fault is that the carried bundle and mesh declarations share one store.** The host cannot +tell "this machine raised this for itself before there was a mesh" from "the mesh told this +machine to have this", so the second overwrites the first completely. + +That is invisible until the two meet, which happens exactly once per mesh: on the first node, +after enrolment, at the moment the control plane first speaks. + +## Why it matters more than a footgun + +**The first node is the only node where the substrate is not the mesh's doing.** Every other node +receives everything it runs from the control plane, so a complete declaration is complete by +construction. The first node raised its own substrate from a file it carried, and the control +plane has never been told about it — so the control plane cannot include it in a declaration even +if it wanted to. + +So the first node is left in a state no other node is in, and the ordinary path destroys it. + +## What is not the answer + +- **Making the control plane send the substrate back.** It does not know what the bundle contained + and should not: the bundle exists precisely because there was no control plane yet. +- **Making apply stop removing orphans.** Removal by omission is how a declaration says *stop + running this*, and losing it costs the property that a node converges on what it was told rather + than accumulating. +- **Special-casing the first node.** [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) + is explicit that its specialness lasts two commands, and this would extend it for ever. + +## The fix + +**The store records where each resource came from** — `carried` or `declared` — and each origin +removes only its own. A declaration removes what the mesh previously declared and never what the +bundle raised; reconciling the bundle removes what the bundle previously raised and never what the +mesh assigned. + +State written before the field existed reads as `carried`, because everything a host had applied +at that point came from its bundle — there was no other way to tell it anything. Guessing the +other way would have the first upgrade remove the substrate, which is this fault arriving through +the change that fixes it. + +**Verified on the scenario that caused it.** A first node raised eleven resources, enrolled, and +was sent the same two-resource declaration. Both applied; the store, the broker and the control +plane were still running afterwards. A second declaration dropping one resource removed that +resource and nothing else, so removal by omission still works — which is the property that had to +survive the fix. + +**What remains open** is what happens when the mesh eventually declares the substrate, which it +must, or the substrate can never be upgraded. Two sources claiming one container is the ambiguity +this issue is made of, narrowed rather than removed: it can no longer happen by accident, and +nothing yet says what it means when it happens on purpose. + +## How it was found + +In the lab, on a sealed machine, by doing the ordinary thing: raise a first node, enrol it, and +tell it something. It was not a test of this — it was the first end-to-end run of the link, and +this fell out of it. + +The declaration was two lines and destroyed a working mesh in under a second, which is worth +holding on to: this is not an edge case reached by trying, it is the first thing that happens. + + +## Two more faults found while fixing it + +Both of the same shape, and worth recording because the shape is the point. + +**A report published to an unbound routing key vanishes.** The control plane bound `enrol` and not +`report`, so nodes announced what they had applied into a void — the broker accepted each message, +found no queue for it, and dropped it. The publisher was told nothing. Reports are now published +`mandatory`, so anything unroutable comes back and is said out loud, and the binding covers every +key a node may publish. + +**A publish failure was being swallowed.** `publishReport` discarded its error, so a node that +could not tell the mesh what it had done looked identical to one that had. That is the fault this +repository keeps cataloguing, written by hand into the newest code in it. diff --git a/04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md b/04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md new file mode 100644 index 0000000..e172319 --- /dev/null +++ b/04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md @@ -0,0 +1,85 @@ +--- +status: fixed +opened: 2026-08-30 +located-in: [mesh-host] +fixed-by: mesh-host — apply attempts every resource and reports every failure +amended-design: +--- + +# 011 — One broken module stops every module after it, for ever + +## Symptom + +A machine was assigned a module declaring a package that does not exist. Every later push to that +machine applied **nothing at all**, and kept doing so. + +Found while proving something else. A test assigned a deliberately-impossible module to a machine +to check that the mesh reports a failure — which it does. A later test on the same machine then +failed, and the evidence said why: + +``` +applied 0 and failed: applying "impossible.nothing": + installing a-package-that-does-not-exist: target not found +0 resource(s) were applied before this and remain +``` + +The broker's queues were **empty**, so the declaration had been delivered and read. The machine +simply stopped at the first failing resource and never reached the rest. + +## Why it matters more than one machine + +- **A machine with one bad module and nine good ones runs none of the nine**, and the mesh reports + "failed" without saying that the rest were never attempted. +- **It cannot be recovered by retrying.** Anything that re-pushes to machines that are behind — the + obvious next feature — would retry a permanent failure for ever and make no progress on + everything else. +- **The order is not the operator's.** Which module is "first" is an accident of resolution, so + which nine modules a broken one blocks is unpredictable. + +## What was there, and what it rested on + +The behaviour had a test asserting it: *nothing after the failure ran*. Its comment cites +[ADR 0010](../../02-DECISIONS/0010-delivery.md). + +**That record does not decide this.** What it says is that a failed *job* stops and names its step +while a reconciler retries forever, as an argument about pipelines against reconcilers. It says +nothing about whether one resource failing should prevent the next from being attempted. The +citation was doing more work than the record supports. + +## The fix, and the argument that was on the other side + +**Everything is attempted, and every failure is reported.** + +The case for stopping is that a resource may depend on an earlier one — a service on the file it +reads. That is real, and it survives: such a service fails its own check and is reported. This host +reads back after every write precisely so a thing that did not work is caught rather than assumed, +so attempting it produces *more* information than skipping it. + +**What is unchanged:** a declaration that cannot be parsed is still refused whole, and nothing is +applied. That is a different thing — *this machine could not do it* against *this was never a +declaration* — and they are fixed in different places. + +### And one shape still stops what follows, which the first fix got wrong + +**An action does.** The first version of this fix continued past everything, and the next lab run +failed at the bootstrap: the store did not answer in three minutes and then said *the database +system is shutting down*. Carrying on past the readiness gate had started the broker and the +control plane against a machine that was not ready, and on a small machine that is how a database +still initialising has its memory taken away. + +**An action is the only shape whose purpose is to make something true *before* the next thing needs +it** — which is why it is the only one with a `verify`. The bootstrap is a row of them: the store +answers, then its databases exist, then their schemas, then the broker. Everything else is +independent state: a package that will not install has nothing to do with a file on the other side +of the declaration. + +So the rule is: **a failed action stops what follows; nothing else does.** Both faults are fixed by +it, and the report says which happened — *these things failed* and *these things failed and the +rest was never tried* are different machines. + +## How this is checked + +`internal/apply`, three tests: an apply with a resource that cannot succeed still applies the ones +after it; every failure is counted, not just the first; and a failed **action** stops what follows +and says so. Each confirmed to fail when its behaviour is removed — including the last, which fails +if actions stop being treated as gates *and* if everything is treated as one. diff --git a/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md b/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md new file mode 100644 index 0000000..3b93ba6 --- /dev/null +++ b/04-ISSUES/012-a-scenario-machine-is-too-small-for-four-images/00-report.md @@ -0,0 +1,88 @@ +--- +status: resolved +opened: 2026-08-30 +located-in: [mesh-lab] +fixed-by: mesh-lab — scenario machines stay at 1 GiB, and the scenario now stocks seven images +amended-design: +--- + +# 012 — Raising a scenario machine's memory stops the substrate coming up + +*Renamed after the first diagnosis turned out to be wrong. What that was, and how it was wrong, is +below — it is the more useful half of this report.* + +## Symptom + +Adding a fourth image to the two-machine scenario made the bootstrap fail every time. The store +container was created, and its readiness check then failed for the full three minutes with +**no output at all**: + +``` +failed store-ready (in mesh-store: ... pg_isready ... exit 1): + the action ran without error and its own verify still fails: docker exited 1: +``` + +Empty after the colon. The check runs `pg_isready` inside the container and prints the store's own +last lines when it gives up; producing nothing means **the container was not running**, which is a +different fault from a database that is slow to start. + +Three images: the scenario raises, both machines join, and nine assertions pass. Four: it never +gets past the store. Reverting the fourth image restores it. + +## The first diagnosis was wrong, and this is why it is worth writing down + +**Two things changed at once.** A fourth image was added to the scenario, and — reasoning that a +machine running a database, a broker and the control plane at once is genuinely small — scenario +machines were raised from 1 GiB to 2 GiB. The bootstrap then failed every time, and the fourth +image was blamed. + +**Removing the image did not fix it. Removing the memory increase did.** Nine assertions pass again +with the extra memory reverted, on a scenario with three images. So the cause is the memory change: +three machines at 2 GiB, on a host also running other work, contend enough that the store container +does not come up at all. + +**The lesson is the ordinary one and it still caught me:** two changes went in together, the failure +was attributed to the plausible one, and an issue was written recording the wrong cause. What found +it was reverting to the exact last-known-good state rather than reverting the suspicious change. + +**What remains untested** is whether a fourth image alone is fine. It probably is. Nothing has +measured it, and the honest state of this issue is that the thing it was opened about was never +demonstrated. + +## What was worth keeping + +**A better diagnostic.** The readiness check now prints what it saw before giving up. That is what +showed the output was empty, which is what said *the container is not running* rather than *the +database is slow* — and which will make the next occurrence of this a diagnosis instead of a +retry. + +## What this blocks + +The mesh running its **own artifact store** — a registry as a module — needs a registry image on +the machine so the module can mirror one, which is the fourth image. The module is written and its +manifest is accepted; what has not been proven is a machine assigned it serving artifacts to +another machine. + +## How this will be checked + +A scenario raised with four images comes up and passes the assertions that three do — which is +what was never actually established. Until then the artifact-store test is not in the shared +scenario, with a note saying where it went and why. + +## Resolved + +*2026-08-31.* The condition this report set was *a scenario raised with four images comes up and +passes the assertions that three do*. The scenario now stocks **seven** and has raised cleanly +many times over, with the machines at 1 GiB where the wrong diagnosis had put them at 2. + +So both halves are settled. **The memory increase was the cause** — reverted, and never +reintroduced. **A fourth image was never the problem**, which this report said had not been +demonstrated either way, and now has been: three more were added on top of it, and the artifact +store the issue said was blocked is proven in the shared scenario rather than kept out of it. + +**The diagnostic that came out of it is what remains valuable.** The readiness check prints what it +saw before giving up, which is what turned *the database is slow* into *the container is not +running*. It has since caught a different fault of the same shape — an action succeeding into a +state its own verify rejects +([04-ISSUES/017](../017-an-action-succeeded-into-a-state-its-verify-rejects/00-report.md)) — which +is the argument for keeping a good diagnostic after the incident that prompted it is gone. diff --git a/04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md b/04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md new file mode 100644 index 0000000..479c70b --- /dev/null +++ b/04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md @@ -0,0 +1,56 @@ +--- +status: resolved +opened: 2026-08-31 +located-in: [mesh-control] +fixed-by: mesh-control — what the mesh computes is applied before what the module declared +amended-design: +--- + +# 013 — A file the mesh computes arrives after the service that needs it + +## Symptom + +Everything the control plane computes for a module — a certificate, a sealed credential, a bound +file, a rule set — was placed **after** that module's own resources in the declaration. The host +applies resources in the order it is given and +[does not sort](../../02-DECISIONS/0005-the-node-host.md), so a service or container declared in a +manifest was applied **before** the file it depends on existed. + +On the first apply the service starts against a missing file and fails. The next reconcile finds +the file there and starts it. + +## Why this matters + +**It repairs itself, which is why nothing caught it.** A fault that is gone by the second attempt +is worse than one that persists: what gets remembered is that the thing works, and the failed +first apply is read as a machine that was briefly slow. The mesh reports a failure, then reports +success, and nobody looks again. + +It was also invisible to every test that existed, because none of them combined the two halves. +Modules with computed files declared no service; modules with a service needed no computed file. +The fault lived exactly in the gap between two repositories' assumptions — the control plane +deciding an order, the host promising not to change it — which is the shape this folder exists for. + +Found by reading, while writing the first module that has both: a firewall whose service must +reflect a rule set the mesh computes. + +## Evidence + +`internal/catalogue/declaration.go` built each module's resource list as +`append(module's own, computed...)` in six places — certificate, authority, needs, secrets, grants, +bindings. `internal/apply/apply.go` iterates `d.Resources` in order, and +`internal/declaration/declaration_test.go` states the rule directly: *order is stated, not derived. +The host must not sort.* + +## What was done + +The computed resources are assembled first and the module's own resources follow. Nothing the mesh +computes is derived from a module's resources, so the order is unconditionally right rather than a +heuristic — there is no case where a module's resource must precede a file the mesh made for it. + +Merged after the computed-resources branch, which replaces a module's resources wholesale and +would otherwise discard everything the mesh had made for it. + +*Checked by a module declaring a service that reflects a rule set, asserting the rule set is first; +and by a module whose resources are computed elsewhere, asserting its credential survives and is +still first — the case the merge point exists for.* diff --git a/04-ISSUES/014-a-key-that-is-present-and-unusable/00-report.md b/04-ISSUES/014-a-key-that-is-present-and-unusable/00-report.md new file mode 100644 index 0000000..e02abb0 --- /dev/null +++ b/04-ISSUES/014-a-key-that-is-present-and-unusable/00-report.md @@ -0,0 +1,48 @@ +--- +status: resolved +opened: 2026-08-31 +located-in: [mesh-host] +fixed-by: mesh-host — a node's serving key is stored in the format a server reads +amended-design: +--- + +# 014 — A node's serving key was present, correct, and unusable + +## Symptom + +A node generates the key it serves TLS with, the mesh certifies the public half, and the +certificate arrives on the machine as an ordinary file. Everything about that worked. But the host +stored the private half in its own encoding — base64 of the raw key — and **nothing that serves +TLS can read it**: not a web server's `ssl_certificate_key`, not Go's `LoadX509KeyPair`, not +`openssl s_server -key`. + +The file was there, owned by root, mode 0600, holding the right key. The certificate beside it was +valid and chained to the mesh's authority. The server would not start. + +## Why this matters + +**Every check that reads the file passes.** The key exists, the certificate exists, the mesh +recorded the public half, the machine reports the declaration applied. The failure surfaces only +when something connects — the worst place to find out, and the place the certificate work was +specifically designed to move away from. + +It is the same shape as [013](../013-a-file-arrives-after-the-service-that-needs-it/00-report.md) +and worth naming as a class: **two halves of one mechanism designed separately, each correct +about its own half.** The control plane issues PEM because that is what a certificate is. The host +stored the key in whatever was convenient, because nothing in the host reads it back — the whole +point of the file is that *something else* does, and that something else was not in view. + +**The generalisation:** where a file exists so a third party can read it, the format is not an +implementation detail of whoever writes it. It is the interface, and it needs a check that reads +it the way that third party will. + +## What was done + +PKCS#8 PEM, which is what every TLS server reads. A key in the old encoding is refused **by name** +rather than reported as corrupt — it is intact, and the remedy is to enrol again, which is a +different action from repairing a damaged file. + +*Checked by writing a key, decoding the file as PEM, parsing it as PKCS#8, and asserting it is the +same key — and, in the lab, by a real handshake from a second machine that verifies against the +mesh's authority and nothing else. A key that parses is not a key a server can use, which is why +the lab check connects.* diff --git a/04-ISSUES/015-a-command-with-no-answer-was-read-as-success/00-report.md b/04-ISSUES/015-a-command-with-no-answer-was-read-as-success/00-report.md new file mode 100644 index 0000000..d1d89cf --- /dev/null +++ b/04-ISSUES/015-a-command-with-no-answer-was-read-as-success/00-report.md @@ -0,0 +1,44 @@ +--- +status: resolved +opened: 2026-08-31 +located-in: [mesh-lab] +fixed-by: mesh-lab — a command with no marker is a failure, not a success +amended-design: +--- + +# 015 — A command that said nothing was read as having succeeded + +## Symptom + +The end-to-end harness runs a command on a machine and reads its exit status from a marker it +appends to the output. When the marker was absent, the parse produced `Number("")`, which is `0`, +and **the command was reported as having succeeded.** + +The marker went missing whenever a command contained a heredoc. Everything was wrapped on a single +line — ` 2>&1; echo "__exit=$?"` — so a heredoc's terminator line became +`MARKER 2>&1; echo "__exit=$?"`, matched nothing, and the heredoc consumed the rest of the script, +the marker included. + +## Why this matters + +**This is the harness lying in the one direction a harness must never lie.** Everything else in +this repository is arranged around the principle that absence must never be indistinguishable from +success — the host says so about a service that does not exist, the builder says so about a build +that failed, the control plane says so about an empty list. The thing that checks all of that had +the fault itself. + +Its reach is every heredoc in the suite, which is how large files are written to machines: a +substrate bundle, a certificate authority, a listener script. Each was written with trailing +junk from the swallowed wrapper, each reported success, and each happened to be tolerated by +whatever read it — until one was a Python script, which did not run, and the test failed on its own +setup. **That failure read exactly like the thing being tested working.** + +## What was done + +`exec 2>&1` on its own first line, so nothing is appended to the command's last line and a heredoc +terminates where it says it does. A missing marker is now a failure, returning whatever was said +so the reason is visible rather than inferred. + +*Checked by the firewall test, whose listener is written with a heredoc: it could not have started +before this, and the assertion that it is reachable before any rule set exists is what makes the +rest of that test mean anything.* diff --git a/04-ISSUES/016-something-after-the-declaration-was-ignored/00-report.md b/04-ISSUES/016-something-after-the-declaration-was-ignored/00-report.md new file mode 100644 index 0000000..a345431 --- /dev/null +++ b/04-ISSUES/016-something-after-the-declaration-was-ignored/00-report.md @@ -0,0 +1,42 @@ +--- +status: resolved +opened: 2026-08-31 +located-in: [mesh-host] +fixed-by: mesh-host — something after the declaration is refused whole +amended-design: +--- + +# 016 — Anything after the declaration in a file was ignored + +## Symptom + +A JSON decoder reads one value and stops. The host's declaration parser used one, so a file +holding a declaration **followed by anything at all** — a stray line, a second declaration, the +tail of a truncated rewrite — parsed as the first value and the rest was never looked at. + +The machine applied something, reported success, and what it applied was not what the file said. + +## Why this matters + +The host refuses a partial declaration everywhere else, in these words: *a host that applied the +parts it understood would leave a machine that looks configured and is not.* This was the same +fault in its quietest form — not a part left out, but a part never seen, with nothing anywhere +saying so. + +**It was live and invisible.** The end-to-end harness had been appending a line to the substrate +bundle by accident ([015](../015-a-command-with-no-answer-was-read-as-success/00-report.md)), and +every bootstrap in every run applied a bundle with a junk line on the end. Nothing failed, so +nothing was looked at, and the corruption was found only by tracing a different bug backwards. + +**That is the argument for refusing rather than tolerating.** A file with something after it is +more likely to be damaged than deliberate: an interrupted write, two files concatenated, a +generator that emitted twice. Applying the first value is applying something nobody wrote. + +## What was done + +After decoding, the parser requires end of input. Trailing whitespace is not "something after it" +— refusing that would make every file an editor writes unusable. + +*Checked by a valid declaration with a stray line after it, with a second declaration after it, +and with garbage after it, each refused; and by the same declaration with trailing blank lines, +accepted.* diff --git a/04-ISSUES/017-an-action-succeeded-into-a-state-its-verify-rejects/00-report.md b/04-ISSUES/017-an-action-succeeded-into-a-state-its-verify-rejects/00-report.md new file mode 100644 index 0000000..8797f41 --- /dev/null +++ b/04-ISSUES/017-an-action-succeeded-into-a-state-its-verify-rejects/00-report.md @@ -0,0 +1,44 @@ +--- +status: resolved +opened: 2026-08-31 +located-in: [mesh-host] +fixed-by: mesh-host — the store's readiness is checked over TCP, not the socket +amended-design: +--- + +# 017 — An action succeeded into a state its own verify rejects + +## Symptom + +The substrate's `store-ready` action waits for the store to answer, then the host runs the +action's `verify` to read back that it worked. Intermittently the host reported: + +> the action ran without error and its own verify still fails + +Both statements were true. The action waited on the **unix socket**; its verify checked the same +way, a moment later, and found nothing. Roughly one bootstrap in three. + +## Why this matters + +**The action and its verify were asking different questions without appearing to.** While the +store initialises it runs a temporary server on the socket only, then stops it and starts the real +one. The action's loop saw the temporary server and exited happy; the verify landed in the gap +between the two. + +So an action can **succeed into a state its own verify rejects** — and when it does, the host's +report is accurate and useless. It says the command worked and the read-back did not, which is +exactly what the mechanism is for, and names nothing a person can act on. It read as a slow +machine, and the remedy people reach for is a longer timeout, which cannot help. + +**The general rule, which the host's design should carry:** an action's verify is the *definition* +of what the action is for. If the action's own waiting decides it is done by a different test than +the verify uses, the two can disagree — and the disagreement surfaces as an intermittent failure +in the one place designed to catch silent success. + +## What was done + +Both check the store over TCP, which the init phase deliberately does not open — so neither can +mistake the temporary server for the real one, and neither can be satisfied while the other is not. + +*Checked by the bootstrap itself, which is where it failed: this action gates everything after it, +so a mesh coming up at all is the check.* diff --git a/04-ISSUES/018-a-provider-on-the-same-machine-was-never-announced/00-report.md b/04-ISSUES/018-a-provider-on-the-same-machine-was-never-announced/00-report.md new file mode 100644 index 0000000..ab638f0 --- /dev/null +++ b/04-ISSUES/018-a-provider-on-the-same-machine-was-never-announced/00-report.md @@ -0,0 +1,51 @@ +--- +status: resolved +opened: 2026-08-31 +located-in: [mesh-control] +fixed-by: mesh-control — something answered on this machine is still bound +amended-design: +--- + +# 018 — A provider on the same machine was never announced to its consumer + +## Symptom + +A module that `binds` a provision receives a file naming where the provider is and what it said a +consumer must know. When the provision turned out to be answered by another module **on the same +machine**, no file was written at all. + +A build machine sharing a node with the registry it pushes to therefore started, connected to the +broker, and looped: *cannot read what the mesh said about the artifact store: no such file or +directory.* + +## Why this matters + +It was deliberate, and the reasoning is in the code: *a file saying "it is on this node" would be a +fact nobody needs and one more thing to keep true.* That is **right about the location and wrong +about everything beside it.** A binding also carries the provider's `serves` block — the port — +and a consumer cannot invent that whether the provider is next door or on the same disk. + +**The failure names nothing.** Every part a person would check was correct: the module resolved, +the machine applied it, the container ran, the credential was delivered and worked. The one file +that did not exist was one the module never asked for by name — it asked for a *provision*, and +the mesh silently decided the answer needed no writing down. The error is about a path, and the +cause is a decision three layers away. + +**And it only appears when two modules land on one node**, which is the ordinary case in a small +mesh and the rare case in a large one. It would have been found in production. + +## What was done + +The binding is written for a local provider too, **when the provider said something a consumer +must know**. That keeps the original intent exactly where it was right: a shell is answered here +and there is genuinely nothing to say about it; a registry is answered here and the port is still +unguessable. + +The address is this machine's own name on the private network, or loopback when it has none — a +machine off the network still reaches itself, and a name nothing resolves is worse than an address +that always works. + +*Checked by resolving a node holding both a provider and its consumer and asserting the binding +carries the port and the address; by the same on a machine with no private network, asserting +loopback; and by the pre-existing check that a provision with nothing to say still writes nothing, +which is the half that was right.* diff --git a/04-ISSUES/019-a-comment-asserting-a-fact-about-a-machine/00-report.md b/04-ISSUES/019-a-comment-asserting-a-fact-about-a-machine/00-report.md new file mode 100644 index 0000000..5ee6ae4 --- /dev/null +++ b/04-ISSUES/019-a-comment-asserting-a-fact-about-a-machine/00-report.md @@ -0,0 +1,59 @@ +--- +status: resolved +opened: 2026-08-31 +located-in: [mesh-control] +fixed-by: mesh-control — the resolver module's claims about machines are checked on a machine +amended-design: +--- + +# 019 — A comment asserting a fact about a machine, which nothing checked + +## Symptom + +The resolver module carried two statements about the machine it runs on. Both read as reasoned, +both were in prose beside the setting they justified, and **both were wrong**: + +| it said | the machine said | +|---|---| +| `127.0.0.54` is free — "not `.53`, that is systemd-resolved's" | systemd-resolved holds **both**; `.54` is its proxy stub. dnsmasq could not create the socket and never started | +| it takes only `127.0.0.55` | listening on a loopback address takes the rest of loopback with it, `127.0.0.1` included | + +A third statement in the same file was true and incomplete in a way that mattered as much: the +config read `/etc/resolv.conf` for upstreams without saying so, and the module that points a +machine at the mesh writes *this resolver's own address* into that file. So its upstream was +itself. Its receive queue filled with 15KB of queries and every lookup on the machine hung. + +## Why this matters + +**The module had unit tests, and they all passed.** They checked that it names an address, that it +reads what the mesh writes, that it restarts when that changes, and that the two asking modules +point where it answers. Every one of those was true while the daemon could not start at all. + +That is not a gap in those tests. It is what a unit test *is*: it confirms the assertion was made, +never that it is true of any machine. **Only a machine knows which of its addresses are spare, or +what a daemon does with a file when it starts.** + +This is [04-ISSUES/003](../003-firewall-scope-is-read-by-no-code/00-report.md) in prose rather +than in a manifest key. There, five manifests carried a `scope:` that read as a restriction and +restricted nothing. Here, a comment read as a reasoned choice of address and chose a taken one. In +both cases *an unenforced rule is indistinguishable from a wrong one, and costs more, because +people believe it* — and a comment is the least enforced rule there is. + +**It cost three full lab cycles**, at fifteen minutes each, because each one revealed exactly one +of the three faults. + +## What was done + +**The wrong statements are corrected, and the correction says what it now knows rather than +asserting a new comfort.** `.55` is written down as *a convention, not a reservation*: if a future +systemd takes it, one line changes. What the module takes is what its claim already said — the +machine's DNS port — rather than a promise about one address. + +**The unit tests hold what a machine has told us.** They assert the module does not take `.53`, +`.54` or `127.0.0.1`, and that it does not read resolv.conf for upstreams. A unit test cannot +discover those facts; it can refuse to forget them. + +**And the order changed.** A module that asserts something about machines is proven on a machine +*before* its assertions are believed — the lab test written first, not last. Written here because +the cost of the old order is measurable: three cycles, forty-five minutes, for a module whose +mesh-side half was correct from the start. diff --git a/04-ISSUES/020-a-certificate-is-issued-and-never-collected/00-report.md b/04-ISSUES/020-a-certificate-is-issued-and-never-collected/00-report.md new file mode 100644 index 0000000..98c3fd8 --- /dev/null +++ b/04-ISSUES/020-a-certificate-is-issued-and-never-collected/00-report.md @@ -0,0 +1,94 @@ +--- +status: open +opened: 2026-08-31 +located-in: [mesh-control, mesh-lab] +fixed-by: +amended-design: +--- + +# 020 — A certificate is issued and never collected + +## Symptom + +Against a real ACME server in the lab, the proxy orders a certificate for a name the mesh routes, +the challenge is answered, the authority **issues the certificate** — and the proxy never obtains +it. Every TLS handshake then fails, and the order is retried indefinitely. + +The client's error, once per attempt: + +``` +http: TLS handshake error: Post "": unsupported protocol scheme "" +``` + +A POST to an empty URL: the certificate's location, on an order the authority considers valid. + +## What is proven, and it is most of it + +Read from the authority's own log rather than inferred: + +``` +Starting 3 validations +authz … set VALID by completed challenge … +POST /finalize-order/ → Order … is fully authorized. Processing finalization +Issued certificate serial 3ef142939115ee88 +``` + +**The hard half works.** The order is created, the HTTP-01 challenge is answered *at the name being +certified* on port 80 through the proxy itself, the authorisation goes valid, finalisation is +accepted, and a certificate is issued. Across one run the authority issued **two** certificates and +accepted finalise **three** times — the client reaches issuance every attempt and fails at the same +step after it. + +**And the policy that guards the quota is proven too.** The second assertion in the same file +passes: no certificate is ordered for a name nothing routes, so a scan cannot spend an account's +rate limit. + +## What is not known + +**Whether this happens against a real authority at all.** Everything above is against Pebble, which +exists to be a test server. The failure is in the last hop between one client and one server, and +may say nothing about behaviour against a public authority. + +## Ruled out + +| | | +|---|---| +| the directory | fetched and complete — `newAccount`, `newNonce`, `newOrder`, `revokeCert` all present | +| the authority's API certificate | covers `127.0.0.1`; the bundle is named explicitly and verification is not skipped | +| a hand-written server config | suspected, and wrong. Replacing it with the server's **own** default config, changing only the challenge port, gives the identical error | +| the finalize URL being empty | the authority logs finalisation being accepted | +| the challenge path | the authorisation goes valid | +| the server version | pinned 2.5.0 behaves exactly as `latest`, so the draft profiles extension is not it | + +## Where it might be + +- **The order's `certificate` field is absent when the client reads it.** The client waits for the + order to become valid and only then fetches, so an empty location on a valid order is the + remaining shape. +- ~~**A moving tag was used.**~~ **Ruled out.** `pebble:latest` advertises a draft *profiles* + extension, so a pinned 2.5.0 was tried: **identical failure**. The scenario now pins it anyway, + which it should have from the start. + +## Why this is filed rather than pursued + +**The mesh-side behaviour is proven and the remainder is interop between two libraries.** Continuing +would be several more twenty-minute lab runs against a server that is not the one production uses, +to chase a defect that may not exist there. + +**What the mesh needed to show, it showed**: a name it routes gets a certificate ordered from a +configured authority, and a name it does not route gets nothing. The configuration is right, the +challenge path is right, and issuance happens. + +## What would close it + +- The same scenario against a different ACME implementation — a second server, or a real staging + endpoint from a machine that can reach one. **If it passes there, this is a Pebble interop + detail and the issue closes with that recorded.** +- Or the client's request captured on the wire, showing what the order actually contained when the + location was read. + +## Evidence + +- `mesh-lab test/integration/certificates.test.ts`, scenario `a-public-name` +- One assertion passes (no certificate for an unrouted name); one fails (a routed name is never + served). diff --git a/04-ISSUES/021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md b/04-ISSUES/021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md new file mode 100644 index 0000000..a08c6fa --- /dev/null +++ b/04-ISSUES/021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md @@ -0,0 +1,90 @@ +--- +status: fixed +opened: 2026-09-01 +located-in: [mesh-control] +fixed-by: mesh-control df62bb5 +amended-design: +--- + +# 021 — A consumer on the provider's machine is given no credential + +## Symptom + +A module that requires something answered **on the same machine** resolves cleanly and is given +**no credential at all**. Two modules, zero needs: + +``` +postgres provides postgres-database, grants /var/lib/postgres/grants +keycloak requires postgres-database, secrets /var/lib/keycloak/database.env +→ modules: 2, needs: 0 +``` + +Nothing is refused and nothing is reported. The consumer's `secrets:` path is simply never +written, and whatever reads it fails later, somewhere else. + +## Where it comes from + +The world a node resolves against is **every other node**: + +```go +for _, n := range nodes { + if n.Name == exclude { continue } +``` + +So a provider on the same machine is never a `Provider` in `world.Offered`, never becomes a +`Needed`, and the credential loop — which walks `resolved.Needs` — has nothing to walk. Every step +is individually reasonable and the sum is a silent gap. + +## Why it was not noticed + +**Everything proven so far was cross-machine.** The lab's provisioner scenarios put the consumer on +one node and the provider on another, which is the interesting case for a *mesh* and the rare case +in practice. The first module to want a database on its own machine was the first real one. + +The postgres provisioner even records the assumption in passing — *"Node is empty for a module on +this machine, which is asking for something local and is not this provisioner's business"* — which +reads as a deliberate exclusion of local consumers. + +## Why the assumption is wrong + +It holds for a process on the machine reaching a unix socket, where the operating system can vouch +for who is calling. **It does not hold for containers**, which is how nearly everything runs here: a +module's containers reach a provider's containers over TCP on a shared network, and the database +asks for a password exactly as it would from another machine. + +**The machine is not a trust boundary once both sides are containers.** Treating it as one gives +the most common arrangement — a service and its database on one node — the weakest handling. + +## What it is not + +Not the same as [`020`](../020-a-certificate-is-issued-and-never-collected/00-report.md) or a +provisioner defect. The provisioner never sees these consumers because the mesh never records +them as consumers. + +## What a fix has to keep + +- **A local consumer still appears in the provider's grants**, so its provisioner creates the role + or bucket or client, exactly as for a remote one. +- **The credential is still sealed**, to the one node that is both ends. The mesh holding a + readable secret for local consumers would be a hole opened for convenience. +- **Refusing must stay refusing.** A requirement nothing answers is still refused; this is about a + requirement that *was* answered. + +## Fixed + +A requirement answered on this machine is still a requirement. Resolution now records a need for +it, so a credential is made, the provider is told who asked, and the consumer's file is written — +the same as if the two were on different machines. + +The reasoning that made it a gap is now written where it was assumed: the machine is not a trust +boundary once both ends are containers, and treating it as one gave the commonest arrangement of +all — a service and its database on one node — the weakest handling. + +Two later issues came out of the same mistaken instinct and are worth reading together: +[`022`](../022-one-credential-per-node-per-provision-not-per-module/00-report.md), where the +machine was treated as an *identity* rather than a boundary, and +[`023`](../023-a-consumer-cannot-build-a-connection-string/00-report.md), where the consumer was +given a password and never told the name to present with it. + +*Closed 2026-09-01. The fix landed the same day and this record was left open by oversight — the +code and the tests were in place for hours while the record still said `located`.* diff --git a/04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md b/04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md new file mode 100644 index 0000000..f62a987 --- /dev/null +++ b/04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md @@ -0,0 +1,113 @@ +--- +status: fixed +opened: 2026-09-01 +located-in: [mesh-control] +fixed-by: mesh-control 0af3ea1 +amended-design: +--- + +# 022 — A credential belongs to a node and a provision, so a second consumer refuses + +## Symptom + +A node running more than one module that wants the same provision **cannot be planned at all**: + +``` +anchor has 3 modules asking for "postgres-database" and they would share one +credential: gitea, keycloak, umami +``` + +**And that is only the loud half.** The consuming node does not refuse at all. Three modules +wanting one database produce **one** need: + +``` +modules=3 needs=1 + name=postgres-database from=anchor for=gitea +``` + +So the first module gets a credential, the other two get no file at all, and each starts and fails +to authenticate with nothing anywhere saying why — the shape of +[`021`](../021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md), on a +different axis. The refusal that reads like a decision is on the provider; the silence is on the +consumer. + +The refusal is correct about what it says. They *would* share one credential, and sharing one is +worse than refusing — a login that opens three databases is not three credentials. But the +arrangement being refused is the ordinary one. **The node this mesh exists to take over runs +eight modules against one database server.** + +## Where it comes from + +A credential is keyed by *(provision, consumer node, provider node)*: + +```go +func (i *Inventory) SecretFor(ctx context.Context, name, consumer, provider string) (Secret, error) +``` + +`consumer` is a **node**. Everything downstream inherits that granularity: `Grant.Consumer` is a +node, the grant's file is named after a node, and the provisioner names the role it creates after +one — `role := mark + c.Node`. + +So the refusal in `ContributionsTo` is not a check that found a problem. It is the only honest +thing that function can do, given a key that cannot tell two consumers apart. + +## Why it was not noticed + +**Every scenario so far had one consumer per node.** That is the natural shape of a small test — +a consumer here, a provider there — and it is the shape of every lab scenario written to date. A +node with two modules wanting a database is not an edge case discovered by fuzzing; it is what a +real machine looks like, and nothing had modelled a real machine yet. + +The refusal also reads as deliberate. It names the modules, explains the consequence, and refuses +rather than picking — the house rule everywhere else. It looks like a decision. It is a limit. + +## Why the granularity is wrong + +The same argument that closed [`021`](../021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md). +There, the machine was treated as a trust boundary and containers made that untrue. Here, the +machine is treated as an *identity* — as though "who is asking" is answered by naming a host. + +**Two modules on one node are as separate as two on different nodes.** They run as different +containers, on different networks, with different data. A key that cannot distinguish them means +the mesh cannot express the thing it is for. + +It also silently weakens what the provisioner does. `mesh_` is one role. Had the refusal not +been there, gitea's login would have opened keycloak's database — and nothing anywhere would have +said so, because from the provisioner's side it created exactly what it was asked to create. + +## What a fix has to keep + +- **The refusal, where it is still right.** Two modules wanting one provision must not silently + share a credential. After a fix they do not share one, so there is nothing to refuse — but a + genuine collision must still refuse rather than pick. +- **A credential per consuming module**, sealed to the node that holds it. Both facts are needed: + the module is who it is for, the node is what it is sealed to. +- **The provisioner names what it creates after the module**, so a login is traceable to the thing + using it, and so withdrawing one consumer does not remove another's. +- **Withdrawal still works.** A module unassigned must lose its login while the others keep theirs + — which is precisely what one role per node cannot do. +- **Existing single-consumer nodes keep working**, since that is every scenario that exists. + +## Scope + +This crosses the control plane, the grant file naming, and every provisioner that names something +after `Consumer`. It is not a local fix, and it is the last thing between the current state and a +node that looks like a real one. + +## Fixed + +Needs fan out per consuming module in one place, after the resolution walk. The credential's key +gains the consuming module, the grant file is named after both halves of the consumer, needs are +matched by provision *and* module, and the provisioners name the role and the access key after the +module rather than the machine. The refusal is gone because there is nothing left to refuse. + +Existing credentials are discarded rather than backfilled: they cannot say which module they were +for, and one is remade and delivered to both ends on the next push, so it costs one rotation. + +A guard was added for PostgreSQL's 63-byte identifier limit, which truncates with a notice rather +than an error — two consumers whose role names agree that far would otherwise become one login, +which is this same fault at a length nobody would think to test. + +What it did **not** fix is [`023`](../023-a-consumer-cannot-build-a-connection-string/00-report.md): +a consumer now receives its own password and still cannot build a connection string, because the +user name is the provisioner's invention and the bound values cannot reach a configuration file. diff --git a/04-ISSUES/023-a-consumer-cannot-build-a-connection-string/00-report.md b/04-ISSUES/023-a-consumer-cannot-build-a-connection-string/00-report.md new file mode 100644 index 0000000..c9853f6 --- /dev/null +++ b/04-ISSUES/023-a-consumer-cannot-build-a-connection-string/00-report.md @@ -0,0 +1,111 @@ +--- +status: fixed +opened: 2026-09-01 +located-in: [mesh-control] +fixed-by: mesh-control 122680b +amended-design: +--- + +# 023 — A consumer is given every part of a connection except the two it cannot invent + +## Symptom + +A module that requires a database is now given its password in whatever shape its configuration +needs ([`022`](../022-one-credential-per-node-per-provision-not-per-module/00-report.md) and the +sealed-placeholder work). It still cannot connect, because a password is not a connection. + +What it is given is a **binding**, as JSON: + +``` +provision postgres-database +from the node providing it +at that node's address on the private network +serves what the provider said a consumer must know — the port +``` + +What it needs, to write `KC_DB_URL` or `GITEA__database__USER`, is the host, the port, the +database name and **the user name**. Two of those are missing, for two different reasons. + +## The user name is nobody's to say + +The provisioner invents it — `mesh__` — and nothing else in the mesh knows that +string. The control plane does not record it, the binding does not carry it, and the consumer +cannot derive it without hard-coding another module's naming convention. + +So the one identifier a consumer must present in order to authenticate is the one thing no part +of the mesh will tell it. It works today only because nothing has yet had to write a connection +string; every proof so far stopped at "the credential arrived". + +## The values cannot reach the file that needs them + +The binding is a JSON document. The consumers are containers reading `KEY=value`, or a program +reading a YAML file, or one reading an attribute inside a different JSON document. A sealed secret +can now be placed inside any of those — the module writes the file with a hole in it and the host +fills the hole on the machine. **The bound values have no such route**, so the half of the +connection that is not secret is the half that cannot be delivered. + +This asymmetry is backwards. The secret is the hard case, because the mesh must not be able to +read it. The host and port are ordinary facts the mesh knows in the clear, and they are the ones +stuck in a document nothing can read. + +## Why it was not noticed + +Every provider so far has been reached by a **provisioner**, a program written for the job, which +reads the JSON because it was built to. The first consumers to need a plain configuration file +were the first real applications. The binding was designed for the program and then handed to the +application. + +There is also a stale comment saying a binding *"carries no credential: the mesh has no way to +issue one yet"*. That stopped being true when [`021`](../021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md) +was fixed. + +## What a fix has to settle + +- **Who names the role.** Either the mesh records what the provisioner will create, or the + consumer contributes the name it wants and the provisioner uses it. The second is more in + keeping with the rest — a consumer already contributes the database name it wants — and it + removes an invented convention rather than documenting one. +- **How a bound value reaches a file.** The symmetric answer to the sealed placeholder, and + simpler: these values are not secret, so the control plane can put them in before sending and + the host learns nothing new. +- **That it stays name-agnostic.** The control plane must not learn what a `postgres-database` + is. What the keys mean is agreed by the requirement's name + ([ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md)), so + whatever is added has to work for a bucket and a mail relay without being told about either. + +## Blocked on this + +Keycloak, Gitea, Mailu and MinIO all have manifests that parse and resolve, and none of them can +start. This is what stands between the module set and a running one. + +## Fixed + +Both halves had the same cause: **the mesh knew something and did not say it.** + +**Who a consumer is, said once.** `ConsumerIdentity(node, module)` is one derivation, sent to the +provider in its grant and to the consumer in its binding — so the two agree by construction rather +than by two conventions that happened to match on the day they were written. The provisioners now +use the name they are given and **refuse to invent one** if the mesh says nothing: falling back to +a name of their own would create a login the consumer could never guess, and everything would +report success. They also refuse a name that does not carry the mesh's prefix, because that prefix +is how withdrawal finds what it made. + +**Bound values reach the file that needs them.** `${bound:provision:key}` is the symmetric twin of +the sealed placeholder and simpler: these values are not secret, so the control plane fills them +in before sending, and the host gains no field and learns no format. `at`, `as` and `from` are +true of any provision; every other key comes from what the provider said it *serves*, so the +control plane still learns nothing about what a `postgres-database` is. + +Keycloak and Gitea now produce complete connection strings — asserted from the manifests on disk, +checking that every part is filled, that no placeholder survives as a value, and that the password +is still a hole only the host can close. + +### Also found, and separate + +The lab run that was meant to prove this failed in a way that looked like the fix being wrong: a +rotation test could not authenticate against a real database. The cause was that the suite +rebuilt the control plane's image and not the provisioner's, so a run with an image built that +minute used a provisioner built the day before. Fixed in `mesh-lab`; it is +[`005`](../005-pipeline-test-harness-unbuildable/00-report.md)'s family — a rebuild covering +most of what a run uses is worse than one covering none, because the run that follows it is +believed. diff --git a/04-ISSUES/024-a-lab-run-stalls-before-the-host-is-placed/00-report.md b/04-ISSUES/024-a-lab-run-stalls-before-the-host-is-placed/00-report.md new file mode 100644 index 0000000..8601342 --- /dev/null +++ b/04-ISSUES/024-a-lab-run-stalls-before-the-host-is-placed/00-report.md @@ -0,0 +1,106 @@ +--- +status: fixed +opened: 2026-09-01 +located-in: [mesh-lab] +fixed-by: mesh-lab 3503ad9 +amended-design: +--- + +# 024 — A run stalls before the host is placed, and says nothing while it does + +## Symptom + +The first scenario of the end-to-end suite raises its machines and then stops. Observed twice on +2026-09-01, both times after the rebuild step grew: + +- **Once mid-run**, during the rotation test: the suite had reported thirteen passes, then the + process ended with no summary, no failure and no receipt. +- **Once from the start**: `a bare machine becomes a mesh` ran for **35 minutes** against a + measured 4.5, produced no output at all, and was still running when it was stopped by hand. + +Both machines were `RUNNING` throughout. The second one was interrogated directly: the anchor VM +answered, and had **no `mesh-host` log and no containers** — so the run had not reached placing +the host. It was stuck earlier, in stocking the scenario's registry. + +## What is not the cause + +- **Not memory.** 84 GiB available, no OOM in the kernel log. +- **Not the daemon.** `incus exec` into the stalled machine answered immediately. +- **Not the changes under test.** The credential work is applied after the host is placed, and the + host was never placed. + +## What changed just before — *and it was not the cause* + +The rebuild step had gone from two artifacts to six, and every image is pushed into the scenario's +registry, which looked like where the stall sat. That was written down as a coincidence rather +than a diagnosis, and it is as well: **stocking takes 34 seconds and always did.** Timed directly, +eight images, before anything was changed. + +The suspicion was the ordinary kind — the thing that changed most recently looks guilty — and the +thing that changed had nothing to do with it. + +## Why it matters more than a slow test + +**A stall is indistinguishable from work.** The suite prints nothing between starting a scenario +and finishing its first test, so four and a half minutes and thirty-five look identical from the +outside — and the operator's only recourse is to guess, which is precisely how a workstation was +left unbootable in August by killing a package manager that was working. + +The first stall is worse: the process ended *silently* after thirteen passes. No summary, no +receipt, nothing that says the run was cut short. A run that stops without saying so is a run +somebody may believe. + +## What a fix has to give + +- **Progress while stocking**, so a long step is visibly a long step. Bytes moved, images pushed, + anything that changes. +- **A receipt when a run is cut short**, saying how far it got. `lastrun` already refuses to guess; + what is missing is it being written at all when the process dies mid-run. +- **A stated timeout on stocking**, so a stall ends as a failure with a reason rather than as a + process somebody eventually kills. + +## The cause + +**The registry machine was addressed by hand and every other machine was not.** + +Machines get a systemd-networkd unit with a static `Address=`, so networkd finishes configuring +the link and reports it `configured`. The registry instead ran `ip addr add` inline. An address +put on a link that way leaves networkd still waiting to configure something it was never told +about, so the link sits at `configuring` — and `systemd-networkd-wait-online` has +`TimeoutStartUSec=infinity`. + +So `network-online.target` is never reached, and **everything ordered after it never starts.** On +these machines that is Docker. `docker load` then blocks on a socket whose daemon is queued behind +a target that will never come, and the three bounded timeouts around it — save, push, load — stack +to thirty-five minutes. + +Measured on one scenario, before and after: + +| | before | after | +|---|---|---| +| the registry's link | `configuring` | `configured` | +| `docker.service` | inactive, 5 jobs pending | active, no jobs | +| the raise | never finished | **87.5 s** | + +These machines have **no DHCP by design** — a scenario is a closed address space and the +declaration owns the addresses — so nothing was ever going to complete that wait. + +## Fixed + +- **The registry is addressed the way every other machine is**, through the same helper. +- **Placing an image waits for the container runtime** and refuses after 120s, naming what systemd + is still waiting on. A stall becomes a failure that says why. +- **The end-to-end test passes `onProgress`.** The raise reported every step and the test threw it + away, which is why thirty-five minutes of silence and four minutes of silence looked the same. + +The suite then ran to completion: **23 of 24**, the one failure a check of its own that flagged +`/var/lib/mesh/builder/broker` as a credential because `/` is in the base64 alphabet. Fixed with +it. + +## Also learned, at some cost + +**A redirected log lags.** Node block-buffers stdout when it is a file, so `> run.log` sits +unchanged for minutes while the run is fine. That was read as a stall twice — the second time +immediately after the real fix, where a buffering artifact argues the fix did not work. The +machines answer instantly and are the source of truth. *"I cannot see progress" is not evidence of +no progress.* diff --git a/04-ISSUES/025-a-module-must-pin-a-digest-and-nothing-produces-one/00-report.md b/04-ISSUES/025-a-module-must-pin-a-digest-and-nothing-produces-one/00-report.md new file mode 100644 index 0000000..9f20414 --- /dev/null +++ b/04-ISSUES/025-a-module-must-pin-a-digest-and-nothing-produces-one/00-report.md @@ -0,0 +1,121 @@ +--- +status: located +opened: 2026-09-01 +located-in: [mesh-control, mesh-host] +fixed-by: partly — mesh-control ee3cc1b +amended-design: +--- + +# 025 — A module must pin a digest, and nothing produces one + +## Symptom + +Every image reference in every example module is **sixty-four zeros**: + +``` +gitea@sha256:0000000000000000000000000000000000000000000000000000000000000000 +``` + +Eighteen of them, across five modules. Each one parses, resolves, and composes into a declaration +a host accepts. None of them could ever start: the machine would reach `docker pull` and stop. + +This is why those modules are *written* and not *running*, and it was not visible from any check +because every check passes. + +## Why nothing caught it + +The host validates the **shape** of a reference and nothing else — that it is `name@sha256:` plus +sixty-four hexadecimal characters. Sixty-four zeros satisfies that exactly. + +That check is not wrong. A host cannot verify a digest exists without reaching a registry, and +reaching a registry is precisely what the design refuses to make it do +([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). The host is the last +place that could catch this and the wrong place to try. + +## The actual gap + +**A manifest must carry a digest, and nothing in the system produces one.** + +- Images the mesh builds are fine: the bundle writes the digest down *after* building, which is + the whole reason the bundle exists in that shape. +- Images from anywhere else — a forge, a mail system, a database — have no path at all. Somebody + has to look up what `gitea:1.22` points at today and paste it in, and nothing re-checks it. + +So the design is coherent about *pinning* and silent about *where a pin comes from*. A person +writing a module is asked for something they cannot reasonably produce by hand, and given a +placeholder shape that passes every gate. + +## What a fix has to keep + +- **The host still refuses a tag.** A digest is what makes a declaration exact, and that must not + soften. The fix belongs where a module is added or built, not on the machine. +- **A person writes a tag; the mesh holds a digest.** A manifest in a repository naming + `gitea:1.22` is readable and reviewable; the mesh resolving that against a registry once, and + recording the answer, is what makes it exact. The module table already records this shape for + source repositories — where it came from, the branch followed, the commit read — and an image is + the same question asked of a registry. +- **Re-resolving is a decision, not a side effect.** A tag that moves must not silently change what + a machine runs. Whatever resolves it records both, so *this pin is behind its tag* is a question + the mesh can answer rather than something discovered on a restart. + +## Cheaply, now + +An all-zero digest is a placeholder and never a real image. Refusing it costs three lines and +would have caught all eighteen the day they were written. It does not fix the gap; it stops the +gap being invisible. + +## What this blocks + +Every module that names a third-party image, which is every module that is not the mesh itself. +The forge and the mail system are otherwise ready to run. + +## Half of it is done + +**A placeholder can no longer reach a machine.** The refusal sits where a declaration is composed, +not where a manifest is parsed — a file awaiting a pin is legitimate, and the design already says +so for artifacts the mesh builds. Composing is the last moment before a machine sees it. + +**The examples now pin images that exist.** Twelve third-party digests were resolved against their +registries without pulling anything, which is also the mechanism the rest of this issue needs: +`docker manifest inspect --verbose` answers *what does this tag point at* in about a second. + +Two faults came free, and both had been invisible for the same reason as the digests: the mail +system's seven images named repositories that **do not exist** — it publishes to a different +registry entirely — and one of the seven had been renamed upstream. Nothing that only checks the +shape of a reference could ever have found either. + +## The mechanism already existed, and this issue was wrong about that + +**Corrected 2026-09-01, the same day.** This was filed saying nothing turns a tag into a digest. +That is false, and the answer had been designed and built before any of it was written. + +A module does not name an image at all. It names an **artifact**, and declares where that artifact +comes from: + +``` +build.artifacts: [{ name: "gitea", kind: "upstream", from: "gitea/gitea:1.22" }] +resources: [{ id: "server", type: "container", artifact: "gitea", … }] +``` + +`kind: upstream` means *an image somebody else built, mirrored into the mesh's own registry and +pinned by the digest it lands with*. The builder produces it; the manifest the mesh holds is +derived, with `artifact` replaced by the real reference and the key removed, because the host has +never heard of that word. A resource naming an artifact nothing produced is refused. + +So the two-document split this issue described as the shape of a fix **is the design**, and it +covers both cases it said were unsolved: an image the mesh builds, and an image somebody else +built. Mirroring also removes something worse than a stale pin — every machine needing a route to +a public registry, and a tag a stranger can move. + +**What was actually wrong was the examples.** They hard-coded image references instead of naming +artifacts, so they inherited a problem the design does not have. Pinning twelve of them by hand +was treating the symptom, and left the reference pointing at a public registry rather than the +mesh's own. + +## What is still open + +**The examples should name upstream artifacts** rather than carry hand-pinned digests. That is the +remaining work, and it is a rewrite of five manifests rather than a mechanism to build. + +The refusal added here stays: a placeholder must not reach a machine whatever the reason it is +there. diff --git a/04-ISSUES/026-the-data-directories-are-not-declared/00-report.md b/04-ISSUES/026-the-data-directories-are-not-declared/00-report.md new file mode 100644 index 0000000..9e39f1a --- /dev/null +++ b/04-ISSUES/026-the-data-directories-are-not-declared/00-report.md @@ -0,0 +1,96 @@ +--- +status: located +opened: 2026-09-01 +located-in: [mesh-control] +fixed-by: partly — mesh-control 53eb000, withdrawn in 83c6a2f +amended-design: +--- + +# 026 — The data directories are mounted and never declared + +## Symptom + +Four modules mount **fourteen host paths** that no resource in those modules declares: + +``` +gitea /services/gitea/gitea +postgres /services/postgres/db-data +minio /services/minio/data/data1-1 +mailu eleven more, including the mail spool and the admin database +``` + +Each is a bind mount on a container. None is a `directory` resource. The mesh has never heard of +any of them. + +## What that costs + +**They are created by the container runtime, as root.** A bind mount whose source does not exist +is created for you, owned by root, with whatever mode the runtime picks. So `owner` and `mode` — +which exist precisely so a module can say who its data belongs to — are silently not applied to +the only directories that hold data. + +**The protection that exists for exactly this does not reach them.** A directory the mesh declared +and no longer wants is *kept*, not removed, when it holds anything the mesh did not put there +([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)). That rule is the +answer to *what happens to my data when a module goes away*, and it is written in terms of +declared directories. **An undeclared one is not protected by it, because the mesh does not know +it is there.** + +So the single rule guarding against data loss covers the configuration directories, which are +cheap to lose, and not the data directories, which are the reason the rule exists. + +## Where it came from + +These manifests were written by reading the arrangement being replaced and carrying its +`docker-compose` files across — service, image, ports, volumes, environment — into the new +manifest's container shape. That shape can express all of it, which is what made the +transliteration feel like progress. + +**A container shape that can express a compose file will be filled in like a compose file.** The +mesh's model is larger than that: a directory is a thing the mesh owns, with an owner and a mode +and a rule about what happens when it is no longer wanted. A volume line borrowed from compose +declares none of it, and nothing complains, because a bind mount source is a string. + +## What a fix has to keep + +- **Every host path a container mounts is declared.** If a module wants a directory on the + machine, it says so, with who owns it and what mode — and gets the removal rule with it. +- **The check is mechanical.** A person comparing volumes against declared directories by hand is + the process that produced this. It is a few lines against the manifest and belongs beside the + other manifest checks. +- **Not by inventing directories at apply time.** The host creating what a mount needs would make + the mesh's ownership of a directory depend on which resource mentioned it first, and would put + the same undeclared path back a layer down. + +## Not yet answered + +**Where a module's data should live at all.** These paths were inherited whole from the +arrangement being replaced, which put everything under one directory per service. Whether that is +right here is a separate question, and a bigger one — it decides what a person backs up, and what +survives a module being removed. + +## Half fixed + +**All fourteen are declared**, across the forge, the mail system, the store and the object store — +each mount now resolves to a `directory` or to a file the module already names. + +**Enforcing it was tried and withdrawn**, and the withdrawal is the interesting half. A refusal +for any mount no resource declares refuses the **builder**, which mounts the container runtime's +socket. That socket is not the builder's data. It does not belong to the module, it already exists, +and declaring it as one of the module's own directories would be a lie that the host would act on. + +So the rule is right about data and wrong about everything else, because the manifest cannot +currently say which a path is. Two kinds of mount are spelled identically: + +- **the directory my data lives in** — created if absent, owned by the module, protected by + [ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md) +- **a machine facility I was granted** — a socket, a device; it exists, the machine owns it, and + the module is being given access to it + +`capabilities` is the closest existing thing to the second and names no paths. Inventing a field +to separate them is a design decision, so it is recorded here rather than made to get a check +green. + +**Until then the manifests are right by coincidence**, which is the state this issue was opened +about. What is kept is a check that every real manifest still parses — worth nothing against this +fault, and the reason the next attempt finds out in a second rather than in a fifteen-minute run. diff --git a/04-ISSUES/027-a-container-cannot-follow-a-file/00-report.md b/04-ISSUES/027-a-container-cannot-follow-a-file/00-report.md new file mode 100644 index 0000000..c88562d --- /dev/null +++ b/04-ISSUES/027-a-container-cannot-follow-a-file/00-report.md @@ -0,0 +1,65 @@ +--- +status: located +opened: 2026-09-01 +located-in: [mesh-host, mesh-control] +fixed-by: +amended-design: +--- + +# 027 — A container cannot follow a file, and a rotated credential is the case + +## Symptom + +A service can say `restart-on`: *these files changed, so I must be restarted*. A container cannot. +It is not in the shape, and the host refuses a declaration that tries. + +So a container reading its password from a file keeps the password it started with, for ever. +Nothing reports anything: the file is right, the container is up, every check passes. + +## Why this is the same fault the mechanism exists for + +`restart-on` is written against exactly this, in the host's own words: + +> a running service does not re-read its configuration. Replace the file, find the service already +> running, do nothing, and the machine keeps behaving the way it did before — while every check +> passes, because the file is right and the service is up. + +Every word applies to a container, and more so. **Nearly everything the mesh runs is a container** +— a database, a forge, a mail system — and a credential arrives as a file it reads at start. + +## What it costs, concretely + +**Rotation does not reach a container.** Rotating a credential replaces the file on the machine and +tells the provider to accept the new one. The provider is a program that reconciles, so it takes +the change. The consumer is usually a container, so it does not. The two ends then hold different +passwords, which is the fault the whole design is arranged to prevent +([ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) records it costing +two days). + +The end-to-end test that proves rotation works uses a consumer that reads the file on each +attempt, so it does not meet this. + +## Why it was not noticed + +The gap is invisible from the control plane. A manifest carrying `restart-on` on a container +composes into a declaration without complaint and is refused on the machine, so the only way to +learn is to run one — which is how it was found, after nine of them had shipped across seven +modules. + +## What a fix has to keep + +- **Declared state, not a command.** `restart-on` is deliberately not *restart this*: it says the + running thing must reflect these files, and the host works out that it does not. Whatever + containers get must keep that shape, because the link may not carry an action + ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). +- **Recreate, not restart, where that is the honest verb.** A container's environment is fixed at + creation. If what changed is an `env-file`, restarting the container is not enough — it has to be + made again. That is a different act from a service reload and should not be described as one. +- **It must not fire on every reconcile.** A container that is recreated whenever the host looks at + it is worse than one that never follows the file. + +## The near alternative, and why it is not enough + +A module can avoid the problem by having its program read the file on each use rather than at +start. That works for something written for this mesh, and is not available for a database, a +forge or a mail system — which is the whole population this is about. diff --git a/04-ISSUES/028-two-things-want-one-port-and-nothing-says-so/00-report.md b/04-ISSUES/028-two-things-want-one-port-and-nothing-says-so/00-report.md new file mode 100644 index 0000000..83b4754 --- /dev/null +++ b/04-ISSUES/028-two-things-want-one-port-and-nothing-says-so/00-report.md @@ -0,0 +1,99 @@ +--- +status: fixed +opened: 2026-09-01 +located-in: [mesh-control, mesh-host] +fixed-by: mesh-control 1f5b70a, 41f7c51; mesh-host b91342a +amended-design: 02-DECISIONS/0038-the-mesh-assigns-the-port.md +--- + +# 028 — Two things want one port, and nothing says so until the machine + +## Symptom + +The database module cannot start on a machine that runs the control plane: + +``` +Bind for 127.0.0.1:5432 failed: port is already allocated +``` + +The mesh keeps its own store on that machine, from the bundle, and it holds 5432. The module +publishes 5432 too. Everything up to the machine is content: it resolves, it composes, it is +pushed, and it is applied — the container is simply the one resource that fails. + +## Why nothing catches it + +**The substrate is not a module.** It arrives from the bundle a host carries, before there is a +mesh to ask. So the control plane has never heard of `mesh-store` and does not know it holds a +port. Resolution can compare modules against each other and cannot compare a module against the +thing the mesh is built on. + +**And nothing compares modules against each other either.** A port is exclusive on a machine in +exactly the way a claim is — one seat, one display server, one artifact store — and the mesh has a +mechanism for that, which ports do not use. Two modules both publishing 5432 would meet the same +wall, one machine later. + +## It has been met before, and worked around + +The end-to-end test that exercises a real database publishes `5433:5432` rather than `5432:5432`. +The workaround is right there, inline, with no note saying why — which is how a constraint becomes +folklore. + +## What a fix has to settle + +- **Whether a module should publish to the machine at all.** Consumers reach a provider by the + machine's address and the port it *serves*, so publishing is what makes that true. An alternative + is that they reach it on the module's own network by name, and nothing is published — which + changes what `serves` means and is a larger decision than it looks. +- **Where the substrate's ports are written down.** Whatever compares them needs to know what the + bundle holds. The bundle is a list of pinned references; what those containers bind is not in it. +- **What a refusal should say.** *5432 is held by the mesh's own store on this machine* is a useful + sentence. *Port is already allocated*, arriving from a container runtime three layers down, is + not. + +## Not the same as a firewall rule + +`listens` already says which ports a module accepts on, and filtering is computed from it. That is +about what may reach a port from elsewhere. This is about two things on one machine wanting to own +the same one, which `listens` does not model and could not answer. + +## Answered in principle + +[ADR 0038](../../02-DECISIONS/0038-the-mesh-assigns-the-port.md), proposed the same day: **the mesh +assigns the machine-side port and a module does not care.** A module cannot choose well, because it +is written once and assigned anywhere — any number it picks is a guess about a machine it has never +seen. + +A port fixed by its protocol — mail on 25, submission on 587 — becomes a **claim**, which is the +mechanism the mesh already has for what is singular on a machine. Two modules wanting 25 is the +same shape as two wanting the seat, and earns the same refusal at assignment rather than at apply. + +The record also names what this issue missed: the same number is written **three times** in every +module — once for the rule set, once for what a consumer is told, once for what the runtime +publishes — and nothing checks that they agree. A module whose `serves` and whose container +disagreed would hand every consumer a port that answers nothing. + +## Fixed + +**The mesh assigns the machine-side port**, from a high unprivileged range, recorded per machine +and module and kept once chosen. A module says the port its software uses, once, in `listens`. The +container's mapping, the rule set, and what a consumer is told are all derived from the assignment +— so the three copies that agreed only because one person wrote them are now one fact. + +**A port the protocol fixes says so**, and is then a claim: one holder per machine, and the second +refused by name at assignment rather than by a container runtime at apply. + +**And the machine says what it already holds.** This was the half that made the issue: the +substrate is not a module, so nothing in the mesh had heard of the store or the broker. The host +already distinguished what it carried from what the mesh sent — that distinction exists so the two +never remove each other — and now records what each resource binds and reports the carried ones. +The allocator treats those as taken. + +What the declaration binds, not what is open: a machine's open ports are a moving target, and +assigning around those would mean a port that was free when it was asked for and taken when it was +used. + +## What it does not settle + +The question underneath, unchanged: **whether a module should publish to the machine at all**. +Assignment makes publishing safe without making it necessary, and consumers reaching a provider on +the module's own network by name would make the question moot for anything inside the mesh. diff --git a/04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md b/04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md new file mode 100644 index 0000000..b4ba944 --- /dev/null +++ b/04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md @@ -0,0 +1,104 @@ +--- +status: fixed +opened: 2026-09-01 +located-in: [mesh-control] +fixed-by: mesh-control be62f49; mesh-lab f85dbb0 +amended-design: +--- + +# 029 — The artifact store cannot be delivered by the artifact store + +## Symptom + +A mesh that has just bootstrapped cannot install a registry. The module describing one resolves, +composes and pushes; the build never completes, because there is nowhere to put what it builds. + +It has never been seen, because the lab always has a registry standing before the mesh asks for +one, and so does any mesh built on a machine that already had one. + +## What it actually blocks + +Not "a registry cannot be installed" — **a mesh cannot hold its own modules.** + +The artifact store is one store for everything a module ships: images, and archives, which are +directories from a module's repository packed and pushed as content-addressed blobs. So it is the +mesh's module catalogue in artefact form, and the same role a separate object store plays in the +arrangement being replaced. + +Until it exists, a module can be described and resolved but nothing it carries can be kept +anywhere. A mesh that has just bootstrapped can therefore run only what its bundle already holds. + +## The cycle + +Three facts, each correct on its own: + +- **A registry module's image is mirrored in.** `kind: upstream` pulls the reference the module + names and pushes it under a name of the mesh's own, so what a machine fetches is pinned by a + digest this mesh assigned rather than by a tag somebody else can move. +- **The builder publishes to the artifact store**, learned from its `artifact-store` binding — + the same binding any consumer of any provision gets. +- **The builder refuses to run without one**: *"a built artifact nobody can fetch is not built."* + +So installing the thing that provides `artifact-store` requires something that provides +`artifact-store`. + +## Why the design already answers it + +The substrate record asks of each candidate *can it grant itself the thing it provides?* The store +cannot create its own database; the broker cannot create its own virtual host; and the registry +**cannot grant itself a repository**. That is why the registry is substrate by role. + +The same sentence answers this. A module that provides the artifact store cannot be delivered +through the artifact store, so its image is not the mesh's to mirror — it is named directly and +pulled from upstream exactly once, which is what the bundle already does for the three images a +first node starts from. + +## The fix + +**The registry module names its image, and is never built.** A container naming +`registry@sha256:…` needs no builder, no binding and no store. Every module after it mirrors +normally, into the registry now running. + +Nothing new is required: naming an image directly is what most modules do. + +## What it costs, and what it does not + +The machine running the first registry needs to reach a public registry once, to pull that one +image by digest. That is already true of a first node, which fetches three images the same way +before a mesh exists. + +It does not weaken pinning. A digest is exact wherever it came from; what mirroring adds is that +the mesh keeps its own copy and does not depend on a tag somebody else controls. For one image, on +one machine, once, the bundle already accepts that trade and says why. + +## The constraint this puts on a registry module + +**A module providing `artifact-store` may not build artifacts of its own** — not its image, and +not a user interface or a tool server shipped beside it. There is nowhere to put them until it is +running. + +A registry that wants more than the upstream image is therefore two modules: one that provides the +store and only names an image, and an ordinary module beside it that builds whatever else and +mirrors it in the normal way. That is a real limit on how a registry module can be written, and it +should be said out loud rather than discovered. + +## How it is checked + +A manifest that both provides `artifact-store` and declares built artifacts is refused where it is +written, naming the cycle. Otherwise the fault surfaces as a build that never returns, on a mesh +too new to have anybody watching it. + +## Fixed + +**The registry module names its image and is added, not built.** The lab's artifact-store test now +walks the only path open to a real first mesh — the manifest goes in directly, no builder and no +store involved — and passes. + +**And the cycle is refused where it is written.** A manifest that provides `artifact-store` and +also declares built artifacts is refused at parse, naming the cycle: building publishes to the +store, so it asks the mesh to put an artifact into the thing that artifact is needed to create. +The provision name became a constant for the rule to turn on. + +What surfaced it in the lab is worth keeping: the test had always *built* the registry module and +passed, because the scenario's stand-in for the public registry was standing there to receive the +push. A prop that quietly covers for the thing under test is how a bootstrap hole stays invisible. diff --git a/04-ISSUES/030-asking-what-a-machine-should-be-re-signed-its-certificate/00-report.md b/04-ISSUES/030-asking-what-a-machine-should-be-re-signed-its-certificate/00-report.md new file mode 100644 index 0000000..dedb8a3 --- /dev/null +++ b/04-ISSUES/030-asking-what-a-machine-should-be-re-signed-its-certificate/00-report.md @@ -0,0 +1,56 @@ +--- +status: fixed +opened: 2026-09-01 +located-in: [mesh-control] +fixed-by: mesh-control 38d4e77 +amended-design: +--- + +# 030 — Asking what a machine should be re-signed its certificate + +## Symptom + +Every machine carrying a certificate reported as *waiting* — not running what the mesh would send +it — for ever. Pushed seconds ago, already behind again. Nothing wrong, nothing failed, nothing +quiet; only a comparison that never came out equal. + +It surfaced as the one red test in four consecutive runs, and wore three other faults' clothes +first: a test racing the apply it asserted on, a status command that wrote to the database it was +reading, a machine starved at its default size. Each was real; each was fixed; the symptom stayed. + +## Cause + +The mesh signs a certificate for a machine's internal name as part of composing its declaration — +and signed **anew on every composition**. Same authority, same key, same name, same validity +window; a fresh random serial each time, because that is what signing does. So the declaration +composed to answer *is this machine current* differed from the declaration sent by exactly one +serial number, every time, deterministically. + +The comparison is a digest, so one changed byte is as unequal as a different world. + +## How it was found, which is the lesson + +Not by deduction — deduction produced the three wrong theories above. The suite was run once with +its scenario kept standing, and the standing mesh was asked twice: `plan`, `plan`, diff. Two +answers seconds apart, identical to the byte but for one serial, in the certificate file. The +diff had one line where four theories had none. + +A verdict machine that can be kept and interrogated is worth more than the verdict. + +## The rule it broke, third find of its kind + +**Issued once and kept.** The port had it, the secret had it, the certificate did not — composed +fresh on every asking, by the same code that holds the other two still. And like +[`028`](../028-two-things-want-one-port-and-nothing-says-so/00-report.md)'s ReleasePorts, the +keeping was designed and never wired: the serving-key migration added a column *"and what was +issued for it"*, and nothing wrote it. + +A kept certificate stands while the name, the key and the clock agree. A node that rejoined with +a new key or changed its name gets a fresh signing exactly as if nothing were kept; so does one +whose certificate is into its last tenth of life. + +## Verified + +Live, on the kept mesh, before any suite run: one push with the fixed binary and the machine +settled; the second machine likewise; then the mesh's own sentence — all doing what they were +told, running what the mesh would send them. diff --git a/04-ISSUES/031-a-machine-becomes-each-thing-it-was-told-in-turn/00-report.md b/04-ISSUES/031-a-machine-becomes-each-thing-it-was-told-in-turn/00-report.md new file mode 100644 index 0000000..97ee87c --- /dev/null +++ b/04-ISSUES/031-a-machine-becomes-each-thing-it-was-told-in-turn/00-report.md @@ -0,0 +1,42 @@ +--- +status: open +opened: 2026-09-02 +located-in: [mesh-host] +fixed-by: +amended-design: +--- + +# 031 — A machine becomes each thing it was told, in turn + +## Symptom + +A machine that was pushed several declarations in quick succession applies every one of them, +oldest first, at the better part of a minute each. Under the lab's suite — twenty-odd pushes in +fifteen minutes — the anchor machine ran minutes behind the newest declaration, and a test that +waited for it honestly timed out while the machine was busy becoming things nobody wanted any +more. + +Only visible since caught-up became an equality: each report now names the declaration it applied, +and the reports arriving were about ever-older ones. Before that, the same backlog hid inside +timestamp comparisons that happened to pass. + +## Why it is wrong, and why it is also right + +Each declaration is complete — the whole machine, not a delta — so applying an old one is never +*incorrect*, only wasted: the machine converges to a state the mesh has already moved past, then +does it again. The queue keeps a disconnected machine's instructions safe, which is right. What +is wrong is only the order of consumption: **a machine asked to be five successive things should +become the last one.** + +## The shape of a fix + +On waking with a queue, drain it and apply only the newest declaration; acknowledge the +superseded ones without applying them. Whether a superseded declaration deserves a report — and +what its outcome should be called — is the real question for the link's vocabulary: silence reads +as a machine that ignored an instruction, and "applied" would be a lie. + +## What it costs today + +Nothing on a real mesh at rest: pushes are far apart. It costs the lab about a doubling of one +test's wait, and it will cost a real mesh exactly when things are busiest — a flurry of changes is +when a machine can least afford to replay history. diff --git a/04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md b/04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md new file mode 100644 index 0000000..c0c70f0 --- /dev/null +++ b/04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md @@ -0,0 +1,135 @@ +--- +status: resolved +opened: 2026-09-04 +located-in: [mesh-sdk, mesh-catalog] +fixed-by: mesh-sdk src/provisioner rework + redis/postgres/minio/umami adapters (ADR 0048) +amended-design: 0048-a-provider-creates-the-credential-the-mesh-minted.md +--- + +# A provider's provisioner seals with a key the mesh has no way to deliver — and does not need to + +## What was observed + +Building the vertical slice for the module runtime (the module runs its own code as its own +process under its own account), a **provider** module — one that stands up a per-consumer +resource and hands back a credential — was assigned to a node and run as a broker-bound +runtime. The runtime hosts the module's provisioner (the sdk's `runProvisioner`), and the +harness opens by reading a **seal key** from `$MESH_SEAL_KEY`, failing immediately without +one. Every credential it produces for a consumer is sealed to that key with the sdk's +symmetric `seal()` (AES-256-GCM, `mesh-sdk/src/primitives/index.ts`) before being written. + +Nothing in the mesh sets `$MESH_SEAL_KEY`. It is read in exactly two places in the sdk and +set nowhere — no manifest, no control-plane code, no host code. So a provider runtime, as +delivered, aborts at start-up. The slice proved the mechanism only by setting a lab-local key +in the manifest by hand. + +## What a trace of the credential path turned up + +The seal key is not a missing delivery. **The whole symmetric-seal provisioner is orphaned, +and it duplicates — badly — a job the mesh already does.** + +- `runProvisioner` reads request files named `*.grant.json`. **Nothing writes those.** +- It writes sealed credential files named `..credential`. **Nothing reads + those** — not the host, not the control plane. The host reports applied-resource digests + upward and never ships credentials; the control plane has no reference to that filename. +- No consumer ever calls the symmetric `unseal()`. Consumers receive **plaintext**. + +Meanwhile the mesh already carries a provider→consumer credential across nodes, with **no +shared key anywhere**: + +- The control plane mints the password once (`secrets.Make`) and seals it **twice, + asymmetrically** — `ForConsumer` to the consumer node's X25519 public key, `ForProvider` to + the provider node's (`mesh-control/internal/secrets/seal.go`, `mesh-host/internal/identity/ + sealing.go`, NaCl box). +- Each host opens its own copy with its own private key on the machine; the plaintext exists + only for the length of one function call (`mesh-host/internal/apply/apply.go`, the + `${secret:name}` substitution — ADR 0024's "the host is the only thing that ever holds + both"). +- `serves` carries no credential and says so; `receives`/`bound` tell each side *where* its + sealed secret is, never the value. + +The two models also **contradict** each other. The sdk's `seal()` comment says the key is "a +per-node passphrase the host holds"; the host holds no such passphrase — it holds an X25519 +private key, and the control plane's own code refuses a shared symmetric key on principle: +"a key both ends hold is a key the mesh would have to distribute, which is this problem again +one level down" (`secrets/seal.go`). A symmetric `MESH_SEAL_KEY` shared between a provider +node and a consumer node is exactly the thing the mesh was built not to have. + +And the provisioner's model is wrong in a second way: its adapter **generates its own +password** (`generatePassword()`) and creates the resource with it — a different password from +the one the mesh mints and hands the consumer. Even with a seal key delivered, a consumer +would authenticate with the mesh's password against a resource created with the provisioner's. + +## Why it matters beyond this instance + +This is not a four-module problem. The provider contract lives in **one place** — the sdk's +`runProvisioner(resource, adapter)` harness — and every provider is built on it. Four exist +today (redis, postgres, minio, umami); a mesh of any size ends up with many. Whatever the +provisioner harness does, every present and future provider inherits, so the orphaned +symmetric seal is a fault stamped into the interface, not into four adapters. That also sets +the cost of getting it wrong: a contract N providers depend on is N migrations to change +later, which is the argument for settling it deliberately now rather than patching around it. + +As written, each provider carries a provisioner that cannot start (no key), and that, if it +did, would create resources with a password it invented — a *different* password from the one +the mesh minted and handed the consumer — and seal them for a reader that does not exist. The +rule the design states, "a consumer receives a sealed credential and unseals it," is enforced +by nothing: no consumer unseals, and no shared key exists to unseal with. + +## The mesh already does this — confirmed + +The premise the fix rests on is not a hope; it is in the control plane today. For a served +interface, `Inventory.SecretFor` mints one password per (consumer, provider) pair via +`secrets.Make`, sealing it to **both** node keys — `ForConsumer` and `ForProvider`. +`SecretsFrom(provider)` is documented as "every credential a provider node was issued, so it +can be told what to create," and `grantsFor` (plan.go) hands the provider node one `Grant` per +consumer carrying `Sealed: ForProvider`. The provider receives, at the path its `receives` +names, one `Contribution` per consumer: the login to create (`As`, derived by the mesh so both +ends agree — 04-ISSUES/023), the consumer's address (`At`) and requested `Values`, and +`Secret`, the file holding that consumer's password sealed to this provider and unsealed by +its host. Everything the provisioner needs is delivered. It reads the wrong files +(`*.grant.json`, which nothing writes) and invents a password instead of reading the one in +`Secret`. + +## The fix this points to + +A **one-place contract change in the sdk harness**, plus re-pointing today's adapters at it — +not per-provider surgery, and inherited correctly by every provider after them: + +- `runProvisioner` reconciles the mesh-delivered `receives` contributions (not `*.grant.json`): + for each consumer, create the resource under the login `As` with the password read from the + delivered `Secret` file, for its `Values`; withdraw the login when a consumer leaves the file. +- The adapter stops generating a password and stops returning a credential — it is handed the + name and the password and only makes the resource exist. Roughly `create({as, password, + values})` / `remove({as})`, no return. +- `sealKey`, `seal()`, `writeSealedCredential`, `MESH_SEAL_KEY`, and the `.credential` file + leave entirely; the consumer already receives its copy through the mesh's own channel. + +This is proposed as ADR 0048, which defines the corrected provider contract, for ratification. + +## Resolution + +ADR 0048 was accepted and implemented on the branches this issue is fixed by: + +- `mesh-sdk` `src/provisioner/index.ts` now reconciles the mesh's `receives` contributions and, + per consumer, reads the mesh-minted password from the file the host unsealed, calling the + adapter to create the resource under the mesh's login. `$MESH_SEAL_KEY`, the symmetric seal, + `writeSealedCredential`, and the `*.grant.json` / `*.credential` files are gone. The symmetric + `seal()`/`unseal()` primitive had no other caller and was removed. +- The four adapters (redis, postgres, minio, umami) were re-pointed at the new contract — + `create({ as, password, values })` / `remove({ as })`, returning nothing. minio's client gained + a secret-key argument so it sets the mesh's secret rather than generating one. +- Proven in the mesh-lab: `provider-uses-mesh-credential` is green — redis creates the consumer's + login with the password the mesh minted, a client authenticates as that consumer and gets PONG, + with no seal key set anywhere. + +Two things were carved out deliberately, neither blocking: + +- **Data provisions are a separate shape.** umami's `analytics` returns a `siteId` umami + *generates*, not a secret the mesh mints, and a contract that returns nothing cannot hand that + back. ADR 0048 is scoped to credential provisions and says so; the provider→consumer return + path for generated data is left to a separate decision. umami compiles and reconciles under the + new harness; only that return is unaddressed, and it never had the seal-key fault. +- **Teardown beyond "remove the login"** — an object store's leftover data — is each adapter's to + name (minio leaves a non-empty bucket for an operator rather than deleting a consumer's data), + not the harness's. diff --git a/04-ISSUES/033-runtime-config-change-does-not-restart/00-report.md b/04-ISSUES/033-runtime-config-change-does-not-restart/00-report.md new file mode 100644 index 0000000..b7089ab --- /dev/null +++ b/04-ISSUES/033-runtime-config-change-does-not-restart/00-report.md @@ -0,0 +1,67 @@ +--- +status: open +opened: 2026-09-04 +located-in: [] +fixed-by: +amended-design: +--- + +# Changing a module's settings does not restart its runtime — config is stale until recreated + +## What was observed + +Rolling the module runtime out to the catalogue (the runtime that serves a module's tools and +runs its events under the module's own account), each tools+events module receives its +configuration the way the design intends: a mergeable config file the module declares, into +which the assignment's settings are merged. The runtime container mounts that file and reads +it once at start-up, when it builds its API client. + +The design for settings says a config file a module owns can be changed **without editing +it** — a person states an intention, the file is regenerated, and the change takes effect. +The decision that config is the assignment's, not the manifest's, is explicitly so that +configuration can be updated *on the fly* and managed from a dashboard. + +For a runtime delivered as a **container**, that last part does not hold. When settings +change, the control plane re-renders the config file on the node — but the runtime container +is only ever recreated when its **spec** changes, and the spec is image, name, env, ports, +volumes and args. The *content* of a mounted file is not part of it. So the file on disk +updates and the process that already read it keeps the value it read at start-up. The new +configuration does not take effect until something changes the container's spec, or it is +recreated by hand. + +A **service** resource has `restart-on`, which names the resources whose change forces a +restart — exactly this problem, already solved, for units. A **container** resource has no +equivalent field, and the apply path for containers never consults the set of resources that +changed this pass. So the one kind of resource that hosts a module's runtime is the kind that +cannot say "restart me when my config changes." + +The effect is quiet, which is the worst part: setting a value appears to succeed (the file is +correct on disk), and the running tools keep answering with the old configuration, or keep +failing to load because the value that would fix them is present but unread. + +## Why it matters beyond this instance + +Every tools+events module converted to the runtime model now takes its URL and credentials +this way, so this is not one module's quirk — it is the config path for the whole catalogue. +The gap turns the headline promise of the settings design ("change it without editing it, on +the fly") into "change it, then recreate the container by hand," which is the manual step the +design existed to remove. And because the file is genuinely updated, nothing surfaces the +staleness; a dashboard that set the value would report success while the mesh kept doing the +old thing. + +Config set **before** the runtime first starts (settings, then assign, then push) does work — +the file is right when the process reads it. So the gap is specifically about *updates* to an +already-running runtime, which is precisely the case the "on the fly" promise is about. + +## Open questions + +- Should a `container` gain `restart-on`, mirroring the service field, so a module can point + it at its config resource? +- Or should the apply path recreate a container when a file it mounts changed this pass — + making mounted-file content behave like part of the spec, without a new field to declare? +- Should the config file's content (or a hash of it) fold into the container spec, so an + ordinary spec-diff already catches it? That restarts on every change with no new mechanism, + at the cost of a spec that is no longer only the container's own declaration. +- Is a restart even the right primitive for a runtime that could instead watch its config + file and rebuild its clients in place — and if so, is that each module's job or the + runtime host's? diff --git a/04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md b/04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md new file mode 100644 index 0000000..01eec09 --- /dev/null +++ b/04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md @@ -0,0 +1,77 @@ +--- +status: resolved +opened: 2026-09-05 +located-in: [mesh-control, mesh-catalog] +fixed-by: ADR 0049 (a slug for the login) + a shorter minted secret (mesh-control) +amended-design: 0049-a-consumers-identity-fits-the-tightest-backend.md +--- + +# The mesh's derived login does not fit every backend's identity rules — S3 rejects it + +## What was observed + +Proving the provider/consumer contract per backend (ADR 0048), redis and postgres passed: a +consumer authenticated against the provider with the login the mesh derived and the password +the mesh minted. **minio failed**, and not on the credential — on the *name*: + +``` +mc: Unable to add a new service account. The access key is invalid. + (access key length should be between 3 and 20). +``` + +The mesh derives a consumer's login as `mesh__` — here `mesh_anchor_bucketuser`, +22 characters. That is a valid postgres role and a valid redis ACL user, so those providers +create it verbatim. S3 access keys are capped at **20 characters**, so minio refuses to create +the service account under it, and the provisioner retries forever while the consumer, holding +that same too-long access key, could never present it either. + +## Why it matters beyond this instance + +ADR 0048 says a provider creates *exactly* the login the mesh derived, so that the two ends +agree by construction — the mesh hands the same name to the provider (to create) and the +consumer (to present). That only holds if the derived name is one every provider can accept. +It is not: the mesh's `as` is a single format with no knowledge of a backend's identity rules, +and S3's are stricter than a database's. Any provider whose backend constrains identifiers more +tightly than postgres — a length cap, a charset, a required prefix — inherits this, and the +failure lands at provision time, per consumer, as an infinite retry rather than a refusal at +assignment. + +This also shows the seam is real, not cosmetic: `as` is doing two jobs — a stable per-consumer +identity the two ends must agree on, and a literal identifier a specific backend must accept — +and those are not always the same string. + +## The shape of a fix (open, not decided) + +- **Constrain the derivation** so `as` is broadly acceptable — short (≤ 20), a conservative + charset, deterministic. This keeps "the provider creates exactly what the mesh derived" true + everywhere, at the cost of a less legible name, and it is a mesh-wide identity change (every + provider that already created the longer name would see it change). +- **Let a provider map `as` to a backend-valid identifier** it derives the same way on create + and on the consumer's behalf — but the consumer is generic and cannot run minio's mapping, so + this only works if the mapped identifier is *delivered back* to the consumer. That is the + data-provision return path this era keeps meeting (umami's siteId, cloudflare's record) and + does not yet have. +- **Declare the constraint on the interface** (`s3-bucket` states its identifier bounds) and + have the mesh derive within them — the most honest, the most work. + +## Open questions + +- Is `as` meant to be human-legible, or is a short opaque token acceptable — i.e., can the + derivation simply be shortened without anyone minding? +- Do redis/postgres actually want the long name, or did it only survive because they are + permissive? If nothing needs it long, the cheap fix is to cap it. +- Does this fold into the same decision as the data-provision return path, or is it separate? + +## Resolution + +Accepted **ADR 0049** (option E): a module declares an optional short `slug`, and the mesh derives +`mesh__`, bounded by the tightest backend (an S3 access key's 20) and refused at +assignment — naming the slug as the remedy — when it still would not fit. The minio grant e2e proved +it: `bucketuser` declares `slug: bkt`, so its access key `mesh_anchor_bkt` (15) is accepted where +`mesh_anchor_bucketuser` (22) was refused. + +Proving that surfaced a **second S3 length constraint on the same credential** — the secret. The +mesh minted a 43-character password (32 random bytes, base64url), and an S3 secret key is 8–40. Fixed +in `mesh-control` `internal/secrets/seal.go` by minting 30 bytes → exactly 40 characters (240 bits, +ample), which fits S3 and every other backend. Both halves of an S3 credential — the access key +(login) and the secret key (password) — now fit the tightest backend, by the same rule. diff --git a/04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md b/04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md new file mode 100644 index 0000000..8c8bd49 --- /dev/null +++ b/04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md @@ -0,0 +1,48 @@ +--- +status: open +opened: 2026-09-02 +located-in: [] +fixed-by: +amended-design: +--- + +# 035 — Reconciling a seed file wipes what grew in it + +## The symptom, as observed + +Found by review of the catalogue examples (2026-09-02), not by an outage — the outage is the +part the design permits to be silent. + +The cache module declares its access-control file as an ordinary file resource with fixed, +empty content. The program that consumes the file requires it to exist at startup, which is +why the manifest declares it at all. But the same file is the one the provisioner writes +consumer users into, and the one the running program persists ACL changes back to. + +A declaration is complete for what the host owns, and the host reconciles what is declared +([ADR 0010](../../02-DECISIONS/0010-delivery.md)). +So every apply that revisits this resource restores the declared content — empty — behind the +running program. Every consumer credential granted since the last apply is removed, the apply +reports success, and nothing anywhere says a grant vanished. + +## Why it matters beyond the instance + +The manifest needed *the file to exist before first start*, and the only vocabulary available +was *the file has this content, forever*. Those are different intentions, and the gap between +them is generic: any resource that a module seeds and something else then legitimately mutates +— an ACL file, a bootstrap configuration a program rewrites, an htpasswd a provisioner appends +to — has the same two owners and the same silent loss on reconcile. + +It is also the mirror image of the boundary ADR 0010 draws so carefully on the *removal* side: +the host never removes what it did not create, but it happily overwrites what it *did* create, +even when what grew inside since is somebody else's work the mesh asked for. + +## Open questions + +- Is the missing thing a create-once file semantic ("present with this content if absent, + untouched otherwise"), or is the real fault that two owners share one file — and the + provisioner, not the declaration, should own it entirely, with first-start ordering solved + some other way? +- ADR 0010 treats every added resource type as a security artefact. Does a create-once + semantic widen what a compromised control plane can express, or narrow it? +- Are there other seeded-then-mutated files already in the catalogue that this failure is + waiting inside? diff --git a/04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md b/04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md new file mode 100644 index 0000000..071d7df --- /dev/null +++ b/04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md @@ -0,0 +1,51 @@ +--- +status: located +opened: 2026-09-02 +located-in: [mesh-control, mesh-catalog, mesh-host] +fixed-by: 02-DECISIONS/0051-shared-data-is-the-operators.md +amended-design: 02-DECISIONS/0051-shared-data-is-the-operators.md +--- + +# 036 — Six modules own what they must share + +## The symptom, as observed + +Found by review of the catalogue examples (2026-09-02). The media stack is several modules — +a library server, the acquisition managers, a download client and their satellites — and each +of them declares the same library and download directories as its own resources. + +The resolver refuses two modules that declare one path on one node, with no exemption for +identical content and no merge. That rule is right in general: two owners of one path is the +class of fault this repository keeps recording. But sharing those directories on one machine +is the entire point of this stack — the download client and the managers must see the same +downloads, the library server must see the same libraries. So the set, as written, refuses +its own only sensible assignment. + +No test co-resolves any two of them, which is why the manifests pass today. The first machine +to be assigned the stack together is where the refusal would have surfaced. + +## Why it matters beyond the instance + +The manifests can express *a directory I own* and nothing else, so a directory that is the +shared workspace of several modules was written six times as six private ones. The intention +— several modules, one filesystem contract between them — has no vocabulary, and this is not +a media-stack peculiarity: any pipeline of modules handing files to each other on one machine +(an ingest directory, a spool, a drop folder) hits the same wall. + +It is also a fork in the design the catalogue has otherwise avoided: the fix could be a new +owning module the others depend on, a shared-resource concept in the manifest, or a statement +that co-located file handoff is not a thing the mesh supports and these modules are one +module. Each answer changes what a module *is*, which is why this is an issue and not a patch. + +## Open questions + +- Is the unit wrong — is a stack that must share a filesystem one module with several + containers, the way the mail module already is? +- If it stays several modules: does one of them own the directories and the rest require + them, and is *requiring a directory from a neighbour* a provision, a claim, or a third + thing? +- The duplicate-path rule protects against genuinely rivalrous owners. Whatever expresses + sharing must not weaken it for the cases where refusal is the right answer — what + distinguishes the two, machine-checkably? +- The mesh's own rule is that a rule states how it is checked: whichever shape is chosen, + what test co-resolves the stack so this class of refusal is caught before a machine is? diff --git a/04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md b/04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md new file mode 100644 index 0000000..ad36933 --- /dev/null +++ b/04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md @@ -0,0 +1,57 @@ +--- +status: located +opened: 2026-09-05 +located-in: [mesh-control, mesh-host, mesh-catalog] +fixed-by: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md +amended-design: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md +--- + +# 037 — A module cannot run its own code at a lifecycle phase + +## The symptom, as observed + +Found while converting the catalogue (2026-09-05), across several modules at once. A module can +declare *things that exist* — a directory, a file with fixed content, a network, a container — but +it cannot declare *a step that runs* at a defined point in its own lifecycle. Three converted +modules need exactly that and have nowhere to put it: + +- **mosquitto.** Its Dynamic Security plugin will not start unless `dynamic-security.json` already + contains an admin client *before the broker's first start* — the broker loads the plugin at + boot. Seeding it is a run-once step that must happen after the file resource exists and before + the container starts. The vocabulary has no "before first start." +- **The database providers (postgres/mongodb/mssql).** First-boot seeding works today only because + the *image* happens to do it from an env var. Anything the mesh itself must run once against the + server — a schema migration, an extension enable, a health gate before the module is announced + ready — has no home. +- The seed-then-mutate family already recorded in [035](../035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md) + is the same shape seen from the *content* side; this is it seen from the *timing* side. + +## Why it matters beyond the instance + +This is not a defect in a module — it is a **capability the module system does not yet offer.** A +real class of modules needs to run their own code at points in the build/install/run lifecycle: +seed-before-start, migrate, post-start health-gate, pre-remove drain. The declarative resource +model deliberately describes *state*, not *steps*, and that is right for what it covers; the gap is +that some modules genuinely have a step. + +**Prior art, and its warning.** An earlier mesh had exactly this as a feature: event-driven +**hooks** that ran custom code at phases of the build/publish/deploy pipeline. It was powerful and +it was **complex to set up and flaky** — which is the real content of this record. The need is not +in question; the cost of the obvious answer is. Whatever shape this takes must not reproduce that +fragility, or it will be worse than the gap. + +## Open questions + +- Is the right unit narrow — a **run-once / init resource** ("run this once, here, in the + lifecycle") — or general — a **per-phase lifecycle hook** on a module, and if so which phases + (build / publish / install / pre-start / post-start / pre-remove)? +- Where does a hook's code run — in the module's own runtime container under its scoped account + (ADR 0043/0047), so it inherits the same isolation as its tools and events? Or is some of it the + host's, before a container exists? +- How is a step made **idempotent and reconcilable** so a re-apply does not re-run it + destructively — the same discipline the resource model gets for free and a step does not? +- What is the smallest version that unblocks the three modules above without rebuilding the old + flaky hook engine? Is "seed-before-first-start" alone enough for now, with the general case + deferred? +- A rule states how it is checked: whatever shape is chosen, what lab scenario proves a hook runs + exactly once, at the right phase, and converges on re-apply? diff --git a/04-ISSUES/038-a-provider-is-announced-at-a-name-its-port-is-not-bound-to/00-report.md b/04-ISSUES/038-a-provider-is-announced-at-a-name-its-port-is-not-bound-to/00-report.md new file mode 100644 index 0000000..9228171 --- /dev/null +++ b/04-ISSUES/038-a-provider-is-announced-at-a-name-its-port-is-not-bound-to/00-report.md @@ -0,0 +1,91 @@ +--- +status: resolved +opened: 2026-09-09 +located-in: [mesh-control] +fixed-by: mesh-control — a same-node provider is announced at the port it is published on +amended-design: +--- + +# 038 — A provider is announced at a name its port is not bound to + +## Symptom + +A module that provides a `from: mesh` provision (observed with the database provider) is +announced to its consumers, by [issue 018](../018-a-provider-on-the-same-machine-was-never-announced/00-report.md)'s +fix, at the node's private-network name — the binding a consumer reads carries +`at: .internal` and `serves.port: 5432`. + +But the provider's container port is **published bound to loopback** (`127.0.0.1:`), +not to the address `.internal` resolves to. So every consumer dials the announced +`.internal:5432`, which resolves to the node's private-network address, where **nothing is +listening** — the port is open only on `127.0.0.1`. + +Observed on a node hosting the provider and several consumers: + +- The consumer's binding file says `"at": ".internal"`, `"serves": { "port": 5432 }`. +- Inside a consumer container, that name resolves to the node's private-network address. +- A connection test from the node: the private-network address on port 5432 is **CLOSED**; only + `127.0.0.1` (on the assigned host port) is OPEN. +- Consumers that touch the database only lazily serve a landing page and *look* healthy; consumers + that require the database at startup crash-loop — one with "acquisition timeout while waiting for + a new connection", another connecting and then timing out on its first query. +- The provider itself is healthy: a direct client on loopback answers instantly, few connections, + no locks. + +## Why this matters + +The announced address and the actual listener disagree, so the binding is a promise the mesh does +not keep. It is not one module's misconfiguration: it is the port-publishing step choosing a bind +address that does not match the `at` the resolver hands consumers, so it fails the same way for +**every** `from: mesh` provider with an off-node-reachable consumer — and, on a single node, for +same-node consumers too. + +It hides well. The provider is up, the credential is correct, the database exists, a manual client +works — every part a person checks in isolation passes. Only a consumer that must use the provision +before it can serve anything reveals it, and it reveals it as a timeout, which reads as slowness or +load rather than "the address was never listening". A mesh that co-locates a provider with its +consumers (the ordinary small-mesh case) is exactly where it bites. + +It also blocks anything that must *reach* a routed/served name from inside the mesh, not just +application traffic — see the internal-CA validation dependency noted in the connectivity design. + +## Diagnosis + +The symptom's first reading — "published on loopback" — was **partly a red herring**. Two things +were tangled: + +1. **The real, current-code defect is a served-*port* mismatch, not a bind address.** A bare + `ports: ["5432"]` is assigned a host port and published as `"15432:5432"` — no bind IP, so on + **all interfaces**, reachable at the node's private-network address. But the port a consumer is + *told* is only re-derived from the assignment on the **cross-node** path. The **same-node** paths + (the resolver's `servedHere`, and the `here()` fallback) settle their served facts *while + resolving* — before the host port is assigned — so they carry the **declared** port (5432), not + the **assigned** one (15432). A co-located consumer is therefore announced + `.internal:5432` while the provider is published on `.internal:15432`, and dials a + port nothing listens on. Cross-node consumers were always fine, which is why it read as "the + small-mesh case." + +2. **The `127.0.0.1:15432` seen in the running lab was a stale build.** Current code's publish step + binds all interfaces; the running instance was raised from a mesh-control predating the ADR 0038 + publish rewrite. The substrate's own store *is* deliberately `127.0.0.1:5432` (a private store + must not be exposed) — correct, and not this bug. + +## Fixed by + +`mesh-control` branch `fix/same-node-provider-announced-port` (`c147a26`): after the host port is +assigned, same-node needs (and the `here()` fallback) are redirected through the same +provision→module→assigned-port lookup the cross-node path already uses, so a co-located consumer is +announced the port that is actually published. Idempotent (keyed by the declared port). Regression +test `TestASameNodeProviderIsAnnouncedAtThePortItIsPublishedOn` asserts the announced port equals +the published host port for a co-located provider/consumer — the next assertion after 018's (which +only checked a binding file exists); verified failing without the change. + +*Not yet merged, and the running lab is additionally stale — proving it end-to-end there needs +mesh-control rebuilt and the affected consumer containers recreated.* + +## Noted, not taken + +Binding the assigned port to the node's private-network address specifically (rather than all +interfaces) would be defence-in-depth and would make the publish address match `at` by construction +— but it is a larger behavioural change entangled with the unenforced firewall scope +([003](../003-firewall-scope-is-read-by-no-code/00-report.md)), so it is left as an option. diff --git a/AGENTS.md b/AGENTS.md index 4fb10e9..6c1cfbc 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,7 +1,7 @@ # Agent instructions — Novox HQ This repository is the source of truth for Novox's mission, research, design and decisions — -today almost entirely those of **Novox Mesh**, its first product ([ADR 0028](02-DECISIONS/0028-hq-is-company-scoped.md)). Implementation lives in the code repositories (see +today almost entirely those of **Novox Mesh**, its first product ([ADR 0019](02-DECISIONS/0019-how-this-repository-works.md)). Implementation lives in the code repositories (see [`00-META/repos.md`](00-META/repos.md)). Before changing anything here, read the playbooks in diff --git a/README.md b/README.md index d8ee801..a647135 100644 --- a/README.md +++ b/README.md @@ -100,7 +100,7 @@ Answered separately, a repository of its own is the better home: changes. Tying documents to a code branch means they merge on the code's schedule. - **The reviewers are different.** A design argument is not reviewed the way an implementation is, and it should not queue behind a build. -- **The scope is wider than one repository.** [ADR 0015](02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) sends most modules out of the +- **The scope is wider than one repository.** [ADR 0001](02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) sends most modules out of the monorepo entirely. Documentation that governs several repositories cannot live inside one of them. @@ -110,4 +110,4 @@ than a mechanism — which is why every decision is recorded in [`02-DECISIONS`](02-DECISIONS/) as it is taken, and why a document that states a rule should say how the rule is checked. -Recorded as [ADR 0019](02-DECISIONS/0019-hq-is-its-own-repository.md). +Recorded as [ADR 0019](02-DECISIONS/0019-how-this-repository-works.md).