Commit Graph
12 Commits
Author SHA1 Message Date
jschoubben a2495e4d8e ADR 0034 (proposed): a test defends a decision
how-we-build 5 already says that if a document states a rule about the
mesh, it says how the rule is verified — an unenforced rule being
indistinguishable from a wrong one, and costing more because people
believe it. That has never been applied to decisions, and a decision
record states the same kind of claim.

The gap was found by review: the lab reached 2,128 lines with 1,072
untested and no stated rule broken, because there is no testing posture in
how-we-build at all. Every decision the lab embodies was verified by hand
and none of it survives the terminal it ran in — which is
04-ISSUES/005 in miniature, coverage assumed rather than checked.

Rejected a coverage percentage: it measures how much code a test touched,
not whether anything important is defended, and would have been satisfied
by testing the parser harder while the hypervisor integration stayed
unasserted.

Rejected test-driven development as a hard rule, and not because it is
wrong in general. Half this implementation was discovery — that the
hypervisor CLI reads a definition from stdin and hangs, that it assigns a
MAC without recording it, that a stock image's networking flushes a static
address. A test written first against undiscovered behaviour asserts a
guess.

So: structure and logic tested first, behaviour against a real system
tested alongside, mocking the boundary forbidden, and a blocking gate as
the definition of done. A test names the decision it defends, which is what
makes the pairing checkable — a decision without one can be found rather
than noticed.

Stated as proposed rather than adopted: 6 requires review by someone who
is not the proposer. Records 0001-0033 predate it and are not retroactively
invalid, but each should acquire a test or an explicit note that it cannot
have one, and until then the rule is aspirational for them — which is the
state 5 warns about, recorded rather than hidden.
2026-08-24 22:22:19 +02:00
jschoubben eab4598494 ADR 0033: a router is scenery, not a node
ADR 0016 makes a lab node a virtual machine, and its reasoning is
fidelity: a node boots a stock image and runs the real install, so it has
to be a real machine or the thing under test is not the thing that ships.

That reasoning does not reach a router. Nothing under test runs on one, it
holds no identity, the mesh never installs anything on it, and no assertion
is ever made about its internals. It exists so packets behave the way they
behave in the world, which is the definition of scenery.

So a router is a system container. What it must reproduce is kernel
behaviour — translation, connection tracking, filtering, forwarding — and a
container has the same kernel.

Verified before deciding rather than assumed. In a plain unprivileged
container: ip_forward and ipv6 forwarding both settable, nftables
masquerade accepted and listed back, and the conntrack timeouts that
mapping_ttl depends on both writable. No privileged mode, no nesting, no
capability grants.

Rejected letting the hypervisor provide NAT, on a stronger ground than
speed: it makes the lab provide what the declaration is supposed to own,
and it cannot express a mapping that expires, a gateway that refuses to
forward, or policy between siblings. The model would shrink to fit the
tool.

The distinction is now load-bearing and has to stay legible: node means
something under test, scenery means something that makes the test real. If
the mesh ever installs anything on a router, it has become a node and this
record no longer covers it.
2026-08-24 01:28:08 +02:00
jschoubben a253afe020 Scenario lifecycle, and how two scenarios coexist
ADR 0032: a scenario is a closed address space. Every segment materialises
as its own isolated link belonging to one instance, so two scenarios raised
from the same declaration hold the same addresses and never meet. The
declaration keeps its literal addresses and they mean what they say —
allocating from a pool would have made them a fiction, so a scenario
reproducing a specific topology would stop reproducing it.

The constraint that follows shapes everything: the lab never reaches into a
scenario over IP. It talks to machines through the virtualisation layer's
own channel. If it reached them by address, the workstation would need a
route into each scenario, and two carrying the same prefix would give it
two routes to one destination — failing not with an error but by one
scenario's traffic arriving in another.

That also makes reachability an honest question. Can this machine reach
that one is asked from INSIDE, by executing on the first, rather than
probed from a workstation that is not on the network and whose opinion
would be a different question with a misleadingly similar answer.

The lifecycle itself: six verbs, of which raise and destroy are enough to
be useful and the rest are what make repetition cheap. Raising is
convergent rather than incremental, because a lab behaving differently
from the thing it tests teaches the wrong habit.

A failed raise leaves the wreckage standing. Tearing down on failure
destroys the only evidence, which is backwards — a scenario that failed to
raise is more interesting than one that succeeded.

Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the
mesh keeps state spanning nodes, so restoring one machine while its peers
move on produces a mesh that has never existed, and faults found there
would be artefacts of the lab.

Closes the declaration's open question about running several scenarios at
once.
2026-08-23 23:53:00 +02:00
jschoubben a72fea5342 ADR 0031 and the scenario declaration
The lab provides the underlay; the mesh builds the overlay. This is the
boundary that decides whether the lab is worth having: a scenario that
assigns overlay addresses, elects the hub and writes peer configuration
certifies its own work — if the mesh's peering is broken, that scenario
still comes up green. The most valuable thing the lab can test is exactly
the part pre-building would replace.

So a scenario declares what a hosting provider and a home router would
provide: segments, which machine sits where at which address, what NAT is
between them, which ports are forwarded, which machines are detached. It
declares nothing about overlay addresses, hubs, peering, names or
certificates, all of which become outcomes to observe.

The declaration has four parts — segments, machines, place, snapshot — and
the two scenario classes differ only in place. That is what makes one a
strict subset of the other rather than a fork.

Research 004's most important finding becomes a format constraint rather
than a footnote: the routable segment must use RFC 5737 documentation
space, because the mesh decides public versus private by matching the
address, and a private range there makes the hub test as unreachable while
the mesh silently never forms. A segment without behind: is routable, and a
non-documentation address in it should be refused before anything is
raised — ADR 0008 applied to a configuration file, since the failure it
prevents has no error at all.

Four things left open, including the one that matters most: a lab machine
is always privileged, so the user and edge profiles have no scenario that
exercises them.
2026-08-23 22:33:52 +02:00
jschoubben 09489a298c ADR 0030: the repository structure, and the rule that names them
The tiers were settled and the product was named, but the repositories
themselves existed only in a research sketch. That had already caused two
problems.

ADR 0029 makes the lab phase 0 of the migration and could not say where it
lives, because no record named a repository.

And the sketch contradicted an accepted record: it listed mesh-hq while
ADR 0028 had decided novox/hq and explicitly rejected that name. A design
resting on research is resting on something that can change without a
decision. Corrected in the research too.

The naming rule, which both earlier records implied and neither stated: a
repository belonging to a product carries that product's prefix; a
company-scoped one does not. That is why this repository is hq and the
mesh's are mesh-*.

Seven repositories recorded — host, substrate, control, surfaces, sdk, lab,
and this one. The lab gets its own: it ships to nobody, outlives any single
tier, and drives virtualisation on a workstation, which nothing else does.
Inside the host it would couple development tooling to a shipped
component; inside the control plane the bootstrap scenario would depend on
a tier that does not exist when it is needed.

Tier 4 is deliberately not decided. Whether the catalogue is one
repository, one per domain or one per application stays open from ADR 0015
and is blocked on research 005 — how many repositories hold domains cannot
be answered before knowing what the domains are. mesh-catalog appears in
the sketch and is not decided by this record.

The cost is stated rather than glossed: seven release cadences where there
is one, and cross-repository changes that used to be one commit.
2026-08-23 22:04:04 +02:00
jschoubben b4904fec7e The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a
complete mesh — forge, coordinator, cascade, verify. That is unusable for
building the new mesh, because all four are tier 2 and do not exist yet.

And research 009 had the sequence backwards. It placed the lab at phase B
as verification of tiers already built, but tier 0 is the component that
takes over a machine's packages, services and network. It cannot be
developed against a machine anyone needs. The lab has to exist before the
thing it will test.

ADR 0029 splits scenarios into two classes. The bootstrap scenario is
virtual machines, the host binary and a pinned bundle, with the verdict
coming from what the host reports about the state it reconciled. The full
scenario is the designed one. The first is a strict subset of the second —
same virtualisation, same networking, same lifecycle, stopping before a
control plane exists — so the second is reached by addition rather than
rework.

The consequence worth having: raising a node from nothing stops being the
least-exercised path in the system and becomes the inner development loop.

It also settles the runner's two jobs. Scenario lifecycle is needed
immediately, because something must materialise and reset a mesh before
anything can be written against it. Assertion execution waits for the full
scenario.

Corrects a stale claim in the design while amending it: it argued
scenarios were affordable with system containers and would not be with
virtual machines. ADR 0016 superseded that reasoning and the text had not
followed.

Issue 007: the lab's first requirement is installed and unusable. The
virtualisation package is present and explicitly installed; both units are
disabled, the operator is in no group, and the client reports the server
unreachable. Not issue 001 again — that is an install failing while
reporting success. This is an install succeeding when success was not the
point. A package is files; a capability is a running service and an
identity permitted to reach it, and the module model has no vocabulary for
the second.
2026-08-23 21:57:14 +02:00
jschoubben c0ae8dec96 Remove two disclosures, and record Nox as the answer to 006
Found by a full scan before making the repository public, which is the
moment the public rule stops being aspirational.

A module name identified a specific laptop model — hardware inventory,
which is operational detail about one installation rather than a lesson
that travels. Generalised.

ADR 0028 named a forge username in a repository path, which the public
rule forbids, and the sentence had also gone stale: the repository it
described was subsequently verified empty of anything unique and removed.
Rewritten to state what happened without the username. Removing a
disclosure from a record is the same class as fixing a path — the rule
that permits it outranks the one that forbids editing.

Issue 006 gains its proposed direction: Nox works from within this
repository rather than these documents being synced into the knowledge
base. Better on three counts — no copy, so no drift; no fourth knowledge
system, which was the original objection; always current.

But it changes the promise, and the issue says so. ADR 0019 promised these
documents would surface BESIDE everything else in a symptom search. An
agent that must be asked is reachable, not surfacing, and the two differ
in precisely the case the operational memory exists for — someone
debugging an error with no reason to suspect HQ knows anything about it.
The question narrows to whether a symptom search finds this content
without the searcher already suspecting it.
2026-08-23 21:40:53 +02:00
jschoubben 93a1231e00 Retire the HAL name where it points forward
Skills take the hq- prefix: they are HQ process workflows, not mesh
workflows, and HQ is company-scoped now. hq-new-research, hq-graduate,
hq-new-issue, hq-diagnose, hq-amend-design, hq-handoff,
hq-sync-constitution, hq-status.

Forward-looking prose becomes Novox Mesh or simply the mesh — the root
README, AGENTS.md, the 00-META README, the mission's module example, and
one to-be document that addressed 'someone working on HAL'.

Three categories deliberately keep HAL, per ADR 0027:

The monorepo is still called hal on the forge. repos.md, every code: field
and every located-in: field name a repository that exists under that name,
and renaming them in prose would make them false.

The as-is layer and the research that measured it describe the system that
runs, and that system is called HAL. 124 modules, 9 daemons, a dead
containerised node — those are observations, not intentions.

Records 0001-0026 are immutable. A record says what was decided when it
was decided, and no record is edited for a name.

Also repoints ADR 0022's link at the renamed skill — a path fix, which the
immutability rule permits, not a change of meaning.
2026-08-23 21:26:09 +02:00
jschoubben 87f4f29cc6 Novox Mesh, Nox, and HQ becomes company-scoped
ADR 0027 — the product is Novox Mesh, shortened to mesh internally. HAL was
never chosen: it arrived with the dotfiles repository this grew out of, it
is borrowed, and it is borrowed from the canonical untrustworthy machine
intelligence, which is an odd flag for infrastructure trusted with
credentials. Timing is the substance of the decision, not an aside — the
skeleton is not built, so renaming costs a search and replace now and a
migration later.

Nox is an identity of Novox, and specifically the agent of the MESH rather
than of a node. Nodes keep their own identities. Nox addresses them, and a
human mostly talks to Nox — which makes it the concrete form of the
mission's vision: state an intent, and the mesh works out which node holds
the thing. It holds no private channel. The gap this opens is recorded:
ADR 0012 binds every agent to a home node, and a mesh-scoped agent has
none, so the model needs extending.

ADR 0028 — HQ is company-scoped, novox/hq, with the mesh as its first
product. Checked rather than assumed: the company organisation already
holds live projects that the mesh builds and deploys, so they are tenants
rather than peers, and the mesh is the ground they stand on. There is also
company work outside the mesh already, which strengthens the case and means
the eventual split is closer than "some day" — so each document's scope is
fixed now, in a table, making that split mechanical instead of
archaeological. The folders are deliberately not restructured yet.

The skeleton takes the new vocabulary: mesh-host, mesh-substrate,
mesh-control, mesh-surfaces, mesh-catalog. Substrate drops to four services
now that identity is a hosted workload rather than a dependency.

Research 009 opens the migration, with the reframing that lowers its risk:
replace the control plane, do not move the workloads. Their data never
moves, so it is re-declared rather than adopted — which keeps adoption out
of scope, as the lab design requires. Self-hosting is the last phase, or a
failed cutover takes away the means to fix it.
2026-08-23 21:20:35 +02:00
jschoubben 143f8ab2f1 research 005: which modules actually change together
The domain-grouping premise is testable, so it was tested before drawing a
list. Co-change across the full history of the module catalogue, current
modules only, platform namespace excluded.

Nine commits in ten touch exactly one module, and 50 of 89 modules have
never been edited alongside anything. Two clusters exist above that floor.

Reachability holds up: proxy, resolver, firewall and VPN genuinely move
together under one intent, three times in recent history. That is the shape
ADR 0017 describes and the only place the measurement finds it.

The provider cluster does not, and this is the finding worth having. Every
multi-provider commit is a cross-cutting manifest change applied N times —
feature detection, hook conventions, volume binds, network scoping. None is
a change to what a database is. Merging them would not have prevented one
of those commits, and the history already shows the fix that worked:
verify by shape in the SDK rather than copying a script into every module.
Move the concern into the machinery, do not merge the modules carrying it.

ADR 0017 keeps its principle and gains a pointer to this narrowing.
2026-08-23 18:48:28 +02:00
jschoubben 3f6d939930 Every decision is a record; the ledger is gone
papa-hq has no ledger. Its root is AGENTS.md, CLAUDE.md, README.md, every
decision is a numbered record, and its graduation playbook has no path for
an unrecorded decision. hal-hq now matches.

The ledger's 41 entries classified as: 10 restating a record, 11 restating
design docs, 15 describing how this repository works with the reasoning
sitting in a README rather than anywhere citable, 3 small rules with no
home, 2 superseded stubs. Mostly a copy — and a hand-maintained index, the
exact pattern ADR 0022 had just rejected for the decision index on the
grounds it drifted after one addition. Keeping one copy of that while
removing another is not a position. It also collided by name with
02-DECISIONS/ in any directory listing.

Nothing was dropped. Records 0019-0025 give the repository decisions the
reasoning they never had: HQ is its own repository and is public, design
has two layers, work moves through playbooks, status lives in frontmatter,
issues have a front door, the numbering is the flow, HQ is the source of
the constitution. 0026 records the ledger's own removal.

The three orphan rules went to how-we-build, where a rule is enforced and
keeps the incident that earned it — the package rule was genuinely
unwritten anywhere. Two lab decisions stated only in the ledger went into
the lab design. "Deliberately not decided" went to the research effort and
design document each question actually belongs to.

The chronological view the ledger provided is now generated from record
frontmatter, which is what it was for.

The cost, stated in 0026 rather than glossed: a record is more work than a
table row, so the risk is a small decision going unrecorded because nobody
wanted to write a document. how-we-build takes rules cheaply, which is the
mitigation, not a solution.
2026-08-23 18:17:59 +02:00
jschoubben c0b35652d0 The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
2026-08-23 18:05:11 +02:00