Commit Graph
269 Commits
Author SHA1 Message Date
jschoubben 4bf7a35568 Close the record on the lab
Playbook 02 and 04 were followed for the substance — decisions before design,
design before build — and skipped for the bookkeeping. This closes that.

004 graduates. Its one open item was "not yet stood up"; the lab is stood up,
and the substitution the effort turned on is now enforced by the validator
before anything is raised rather than left as a thing to remember. Its
certificate conclusion has a home in 01-end-to-end-testing and is designed but
not built — implementation is a third axis, and an effort graduates on its
conclusions.

One item leaves 004 without a home and is recorded rather than lost: the reverse
proxy does not set caServer, so it defaults to the production endpoint.

The two lab designs read `designed` while running in production of a sort, so
they become `in-progress`.

And the lab gets an as-is document, which it did not have. It records what runs
including the parts nobody would choose again: that `place:` is refused and the
lab therefore raises EMPTY MACHINES, that the drawing shipped with no design
document behind it, that a router is tagged as a machine for a reason found by a
bug, and that the integration suite raises two of five scenarios while both
faults found so far lived in the three it does not.

006 stays active, deliberately. Two of its open questions ARE the tier 0 design
— whether absorbing six concerns makes the host too large, and whether an
unprivileged node earns a place in the inventory. Playbook 04 is explicit that
an open question is a reason to research, not to build around.
2026-08-25 01:55:52 +02:00
jschoubben 8efa063f21 The snapshot question is answered by a test
The lifecycle design asked whether a scenario snapshot needs the machines
stopped. The integration test answered it on its first run: no, but they
must be flushed.

A snapshot captures disk and not memory, so a write still in the guest's
page cache is absent from it — not stale, absent. A file written seconds
before a snapshot did not survive the restore.

Flushing first buys write-durability. It does not buy
application-consistency: anything mid-transaction is still captured
mid-transaction, and that limit is now stated rather than left implied.
2026-08-24 22:26:47 +02:00
jschoubben eab4598494 ADR 0033: a router is scenery, not a node
ADR 0016 makes a lab node a virtual machine, and its reasoning is
fidelity: a node boots a stock image and runs the real install, so it has
to be a real machine or the thing under test is not the thing that ships.

That reasoning does not reach a router. Nothing under test runs on one, it
holds no identity, the mesh never installs anything on it, and no assertion
is ever made about its internals. It exists so packets behave the way they
behave in the world, which is the definition of scenery.

So a router is a system container. What it must reproduce is kernel
behaviour — translation, connection tracking, filtering, forwarding — and a
container has the same kernel.

Verified before deciding rather than assumed. In a plain unprivileged
container: ip_forward and ipv6 forwarding both settable, nftables
masquerade accepted and listed back, and the conntrack timeouts that
mapping_ttl depends on both writable. No privileged mode, no nesting, no
capability grants.

Rejected letting the hypervisor provide NAT, on a stronger ground than
speed: it makes the lab provide what the declaration is supposed to own,
and it cannot express a mapping that expires, a gateway that refuses to
forward, or policy between siblings. The model would shrink to fit the
tool.

The distinction is now load-bearing and has to stay legible: node means
something under test, scenery means something that makes the test real. If
the mesh ever installs anything on a router, it has become a node and this
record no longer covers it.
2026-08-24 01:28:08 +02:00
jschoubben e88b448145 The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing
1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle
falls from ~90s, unbounded at worst, to ~15s dominated by a boot that
cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write
and did not without it.

The consistency matters as much as the speed: three consecutive snapshots
took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never
finished.

One honest counter-observation recorded: launching onto the fresh
copy-on-write pool was slower, 20.2s against 14.3s, because the image had
to be unpacked into a pool that had never seen it. Paid once per pool, and
dwarfed by what snapshotting saves, but it went the other way.

Doing the measurement produced the answer to how the lab installs on a
clean machine, because both failure modes appeared while doing it.

Installed is not available: the daemon was present with units disabled and
no group. Issue 007.

Available is not adequate, and this is worse: with the storage tooling
absent everything worked and snapshots were seventy-six times slower.
Nothing failed, nothing warned. That is a variant the mesh has not
catalogued — its usual failure is reported success and did nothing; this is
reported success and did it seventy-six times slower, which no error
surface catches because nothing is wrong.

So the lab verifies CAPABILITY, never installation, and refuses to run
degraded rather than warning — a warning about a slow inner loop is read
once and ignored forever. Prerequisites may arrive from a mesh module or
from the lab's own bootstrap, and the second path is required rather than
convenient: a lab installable only by a mesh cannot host the development
of the mesh that installs it.

The lab is the second thing installed by hand, after the node host, and for
the same reason: something has to be first, and pretending otherwise
produces a circularity papered over by a script nobody exercises.
2026-08-24 00:14:57 +02:00
jschoubben 98bcd5cc49 Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than
arguable, so it was measured. One virtual machine on a workstation with
hardware virtualisation and NVMe.

Raising: the launch call returns in 3.4s, the machine is actually usable
after 14.3s. The gap is a design constraint — raise must wait for the
second number, because reporting the first would be transport reported as
effect, which is the mesh's own recurring failure.

Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full
copy; nothing is shared. Restore: 10.4s, usable again after 20.1s.

The second snapshot exceeded two minutes and never completed. That is the
more troubling number: snapshot cost here is not merely high, it is
unpredictable, and a loop with a variable multi-minute step is one nobody
trusts.

Projected to a four-machine scenario, a reset-and-rerun cycle is about a
minute and a half at best and unbounded at worst, before any of the mesh's
own work begins. That is too slow for an inner loop, and ADR 0029's whole
argument — that making the bootstrap path the inner loop turns the
least-exercised code into the most-exercised — holds only while resetting
is cheap.

The cause is not virtual machines. Hardware virtualisation is present and
machines boot in fourteen seconds. It is that the daemon offers exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap
snapshot. The btrfs kernel module is available; btrfs-progs is simply not
installed, which is the entire reason the driver is absent.

The copy-on-write comparison was deliberately NOT run, because running it
would mean installing a package by hand — which the rules forbid and which
would have made the measurement unreproducible. So the honest statement is
that the current configuration is too slow and the likely fix is known but
unverified, rather than that btrfs fixes it.
2026-08-24 00:07:22 +02:00
jschoubben a253afe020 Scenario lifecycle, and how two scenarios coexist
ADR 0032: a scenario is a closed address space. Every segment materialises
as its own isolated link belonging to one instance, so two scenarios raised
from the same declaration hold the same addresses and never meet. The
declaration keeps its literal addresses and they mean what they say —
allocating from a pool would have made them a fiction, so a scenario
reproducing a specific topology would stop reproducing it.

The constraint that follows shapes everything: the lab never reaches into a
scenario over IP. It talks to machines through the virtualisation layer's
own channel. If it reached them by address, the workstation would need a
route into each scenario, and two carrying the same prefix would give it
two routes to one destination — failing not with an error but by one
scenario's traffic arriving in another.

That also makes reachability an honest question. Can this machine reach
that one is asked from INSIDE, by executing on the first, rather than
probed from a workstation that is not on the network and whose opinion
would be a different question with a misleadingly similar answer.

The lifecycle itself: six verbs, of which raise and destroy are enough to
be useful and the rest are what make repetition cheap. Raising is
convergent rather than incremental, because a lab behaving differently
from the thing it tests teaches the wrong habit.

A failed raise leaves the wreckage standing. Tearing down on failure
destroys the only evidence, which is backwards — a scenario that failed to
raise is more interesting than one that succeeded.

Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the
mesh keeps state spanning nodes, so restoring one machine while its peers
move on produces a mesh that has never existed, and faults found there
would be artefacts of the lab.

Closes the declaration's open question about running several scenarios at
once.
2026-08-23 23:53:00 +02:00
jschoubben e88a6df924 Public networks are unrelated, and routed rather than bridged
Caught in review: every public address sat in one /24, which made the
three of them look like one network. They are not. The internet is a very
large number of unrelated networks routing to each other, and a machine in
one is many hops from a machine in another with no shared broadcast domain
between them.

Putting them in one prefix would have quietly made four false things true
in the lab: machines resolving each other by ARP and talking directly, TTL
never decrementing, broadcast and multicast crossing between them, and any
two being adjacent.

The third is not hypothetical. The as-is layer records that mesh names are
deliberately not multicast names, after a delay and a one-node-only failure
mode. A lab where the internet is one broadcast domain would let a node
discover a peer by multicast that it could never discover in production,
and report success — the exact false green this effort exists to prevent.

So a scenario has one public segment per public NETWORK, each with its own
unrelated prefix, wired together through a router and never onto a shared
bridge. That is a property of how the lab wires them rather than a field
anyone sets, because no correct scenario has two public networks adjacent.

Addresses now spread across all three RFC 5737 ranges plus RFC 3849 /48s,
chosen to look nothing like each other, and a foreign private network uses
someone else's RFC 1918 range rather than a documentation one.

All three examples in the document rewritten, since two of them still
showed a single flat internet segment and contradicted the new rule.
2026-08-23 23:49:38 +02:00
jschoubben a873088140 A worked example: the whole model applied to an ordinary mesh
The shape research 004 identified — one machine with a routable address,
one publicly named but behind a household connection, one stationary on
that network, one that roams — written out with every field the model has,
in role names and documentation addresses.

It shows the ISP modem doing nothing, because in bridge mode it is a media
converter: it changes the physical medium and leaves the packets alone, so
it creates no IP-level fact and appears nowhere. In router mode it would be
a second gateway and publishing would need a rule on both, which is the one
case the model still cannot express.

It shows two segments sharing one gateway declaration, which means one
gateway machine, and a policy rule between them that is asymmetric because
useful ones almost always are.

And it shows what is deliberately absent. Research 004 recorded overlay
addresses, hub election and names for exactly this topology, and none of
them appear: a scenario must not state what the mesh is responsible for.
Given the declaration, whether a hub is elected, whether the NATed
machine's endpoint is learned, and whether the roaming machine re-forms
after moving are all observed rather than arranged. The absence is the
point.

The run at the end moves one identity through four positions — home,
foreign network, asleep, home again — against a foreign gateway whose
mapping expires in 30 seconds, which is why a number is there rather than
a boolean.
2026-08-23 23:43:09 +02:00
jschoubben b1874f1d0f Segment policy, shared gateways, and what the model leaves out
Asked whether a real setup is coverable — router, modem, access points —
the answer splits, and one part was a genuine gap.

Most equipment is invisible and the omission is deliberate. The test: does
the device change what an IP packet can do? A switch moves frames within a
segment. An access point bridges wireless clients onto one — a machine on
wifi and a machine on cable are the same machine to IP. A controller
configures equipment and has no packets of its own. Modelling any of them
adds a fixture with no fault to catch.

Two entries in that list do matter. A modem in bridge mode is a media
converter and invisible; in router mode it is a second gateway, which is
double NAT — expressible as nested segments, but publishing through two
gateways still is not, and that is now named as the one real absence.

And VLANs are segments, which exposed the gap: inter-segment policy was
inexpressible. inbound: is a HOST firewall, per machine. A segmented
router enforcing rules between networks is a different thing and blocks
traffic regardless of what the destination thinks — a node behind such a
rule cannot be reached even by a peer that knows exactly where it is.

policy: states it as a fact about a pair rather than a property of either,
defaulting to allowed and asymmetric by design, because the useful
configuration is almost always one-directional.

Segments may also share a gateway: identical gateway declarations mean one
gateway machine, not two, because that is what a VLAN-capable router is —
and two routers sharing an address would not work anyway.
2026-08-23 23:32:30 +02:00
jschoubben 274bd3b304 Close the missing axes — and address family changes the model
Address family was not a field. IPv6 usually has no NAT, so a machine
behind a household gateway is typically unforwardable on v4 and DIRECTLY
ATTACHED on v6, at the same moment. The three positions therefore apply
per family, and reachability is a property of (machine, family) rather
than of a machine.

The consequence is bigger than the syntax: 'can these two nodes reach each
other' stops being a yes/no question. It is asked once per family, and the
asymmetric answers are the interesting ones. A mesh treating reachability
as one fact per node reaches a peer over one family, fails over the other,
and reports whichever it tried. That distinction did not exist in the model
and would have been found by a failure rather than by reading.

Two fields follow from it. inbound: allow|deny became necessary because
with NAT unreachability was implied by topology, while a globally routable
v6 address is reachable unless something refuses — so refusing has to be
sayable or v6 addressing silently implies reachability. And nat: became a
list of families rather than a boolean, because a real gateway translates
v4 and routes v6 and a boolean cannot say that.

mapping_ttl closes the keepalive gap: a mesh holding a connection through
NAT without refreshing it works perfectly until the far side goes quiet
for longer than the mapping lives.

segments[].mtu closes the fragmentation gap: an overlay adds a header, so
a tunnel over a reduced-MTU path establishes a connection and then
silently drops large packets.

at: takes a list, so a multi-homed machine is expressible — which the
model already implicitly required, since a border machine sits on two
segments.

v6 uses RFC 3849 documentation space, the exact counterpart of the RFC
5737 rule and load-bearing for the same reason.

Remaining: nested forwarding and an address changing in place, both
extensible when needed. Path quality stays deliberately out — it changes
performance, not correctness, and modelling it makes a network simulator
rather than a fixture.
2026-08-23 22:59:31 +02:00
jschoubben b944904f1a Audit the scenario model for generality, and fix what it found
The question is not whether the model covers our mesh but whether it can
express any mesh. Audited against the axes a deployment varies along, with
the standard being every property that changes how the mesh BEHAVES rather
than every property a network has — bandwidth does not change correctness,
MTU does.

One real bug, now fixed. A segment with no gateway was read as the
internet, which made an isolated network inexpressible: a LAN with no route
out would have been treated as public and forced onto documentation
addresses. Segments now state kind: public or private, and a private
segment with no gateway is an island. A mesh spanning a site with no
internet is a real topology.

One modelling error, now corrected. The three positions were framed by
ownership — a gateway you control versus one you do not. The axis is
forwardability. Carrier-grade NAT is your own connection and is still
unforwardable, so it belongs with the café network. Gateways gain
forwardable:, independent of nat:, and publishing through an unforwardable
one is a declaration error because that is the constraint being reproduced.

Three genuine gaps recorded in priority order. Address family: cidr is
implicitly v4, and a v6-only node is not exotic — a mesh that assumes v4
fails there completely rather than partially, which makes this a second
world rather than a refinement. Expiring NAT mappings: without them
keepalive behaviour is hoped for rather than tested, and for a mesh mostly
behind NAT that is the fault that shows up after an idle night. MTU:
tunnels fragment, and a smaller-MTU path establishes a connection that then
silently drops large packets — the exact shape this effort exists to stop
shipping.

Latency and loss are deliberately out: they change performance, not
correctness, and modelling them makes a network simulator rather than a
fixture.

Also adds a NAT primer, because the three positions are consequences of it
and the document should not assume the reader already knows why a mesh
dials outward and never inward.
2026-08-23 22:48:46 +02:00
jschoubben e65e5809dc The scenario declaration gets a real network model
forwarded: [443] was the tell. It implied a destination-NAT rule while
never saying from which address, and the address is the whole point: a
household's public address is what a peer records as the endpoint when a
machine there dials out, and what a public name for a published machine
there resolves to. It was decoration in the old shape and is load-bearing
in this one.

The model now names three positions a machine can be in, because they are
genuinely different and the mesh has to cope with all three. Directly
attached, with its own routable address. Behind a gateway you control,
reachable only through a forwarded port at the gateway's address. Behind a
gateway you do not control, reachable not at all, with an apparent address
belonging to someone else's router that changes when the machine moves.
The third is the hard one and the one that breaks reachability
assumptions first.

A gateway now carries three facts instead of a boolean: the parent
segment, the address the world sees the network as, and whether addresses
are translated — so a routed range is expressible as well as ordinary
household NAT. published names the gateway it forwards through, which is
how a machine on a LAN that itself has a public address is stated, and
publishing on a foreign gateway is a declaration error because that is
exactly the constraint being reproduced.

Moving a machine between positions becomes a lifecycle operation rather
than a declaration: the same identity at home, then on a foreign network,
then asleep, in one run. Whether the overlay survives that and notices the
endpoint changed is observed, never arranged.

All three RFC 5737 ranges are now allocated a job — the internet segment,
a foreign network, and a spare — with private segments kept byte-identical
to production because those addresses mean the same everywhere.

New open question worth having: a real gateway forgets NAT mappings after
a timeout, and whether a scenario can say so decides whether keepalive
behaviour is testable or merely hoped for.
2026-08-23 22:42:27 +02:00
jschoubben a72fea5342 ADR 0031 and the scenario declaration
The lab provides the underlay; the mesh builds the overlay. This is the
boundary that decides whether the lab is worth having: a scenario that
assigns overlay addresses, elects the hub and writes peer configuration
certifies its own work — if the mesh's peering is broken, that scenario
still comes up green. The most valuable thing the lab can test is exactly
the part pre-building would replace.

So a scenario declares what a hosting provider and a home router would
provide: segments, which machine sits where at which address, what NAT is
between them, which ports are forwarded, which machines are detached. It
declares nothing about overlay addresses, hubs, peering, names or
certificates, all of which become outcomes to observe.

The declaration has four parts — segments, machines, place, snapshot — and
the two scenario classes differ only in place. That is what makes one a
strict subset of the other rather than a fork.

Research 004's most important finding becomes a format constraint rather
than a footnote: the routable segment must use RFC 5737 documentation
space, because the mesh decides public versus private by matching the
address, and a private range there makes the hub test as unreachable while
the mesh silently never forms. A segment without behind: is routable, and a
non-documentation address in it should be refused before anything is
raised — ADR 0008 applied to a configuration file, since the failure it
prevents has no error at all.

Four things left open, including the one that matters most: a lab machine
is always privileged, so the user and edge profiles have no scenario that
exercises them.
2026-08-23 22:33:52 +02:00
jschoubben 4d387d2998 Hand the lab design off to mesh-lab
Playbook 04: the design names its owner and flips to in-progress. code
moves from hal to mesh-lab, and decisions gains 0029 and 0030 — the design
now rests on three records rather than one.

repos.md marks mesh-lab as the one target repository that exists. The
other six remain the target, not the present, and saying so is the point:
a map that lists repositories which do not exist is a map that will be
believed.
2026-08-23 22:26:30 +02:00
jschoubben b4904fec7e The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a
complete mesh — forge, coordinator, cascade, verify. That is unusable for
building the new mesh, because all four are tier 2 and do not exist yet.

And research 009 had the sequence backwards. It placed the lab at phase B
as verification of tiers already built, but tier 0 is the component that
takes over a machine's packages, services and network. It cannot be
developed against a machine anyone needs. The lab has to exist before the
thing it will test.

ADR 0029 splits scenarios into two classes. The bootstrap scenario is
virtual machines, the host binary and a pinned bundle, with the verdict
coming from what the host reports about the state it reconciled. The full
scenario is the designed one. The first is a strict subset of the second —
same virtualisation, same networking, same lifecycle, stopping before a
control plane exists — so the second is reached by addition rather than
rework.

The consequence worth having: raising a node from nothing stops being the
least-exercised path in the system and becomes the inner development loop.

It also settles the runner's two jobs. Scenario lifecycle is needed
immediately, because something must materialise and reset a mesh before
anything can be written against it. Assertion execution waits for the full
scenario.

Corrects a stale claim in the design while amending it: it argued
scenarios were affordable with system containers and would not be with
virtual machines. ADR 0016 superseded that reasoning and the text had not
followed.

Issue 007: the lab's first requirement is installed and unusable. The
virtualisation package is present and explicitly installed; both units are
disabled, the operator is in no group, and the client reports the server
unreachable. Not issue 001 again — that is an install failing while
reporting success. This is an install succeeding when success was not the
point. A package is files; a capability is a running service and an
identity permitted to reach it, and the module model has no vocabulary for
the second.
2026-08-23 21:57:14 +02:00
jschoubben 93a1231e00 Retire the HAL name where it points forward
Skills take the hq- prefix: they are HQ process workflows, not mesh
workflows, and HQ is company-scoped now. hq-new-research, hq-graduate,
hq-new-issue, hq-diagnose, hq-amend-design, hq-handoff,
hq-sync-constitution, hq-status.

Forward-looking prose becomes Novox Mesh or simply the mesh — the root
README, AGENTS.md, the 00-META README, the mission's module example, and
one to-be document that addressed 'someone working on HAL'.

Three categories deliberately keep HAL, per ADR 0027:

The monorepo is still called hal on the forge. repos.md, every code: field
and every located-in: field name a repository that exists under that name,
and renaming them in prose would make them false.

The as-is layer and the research that measured it describe the system that
runs, and that system is called HAL. 124 modules, 9 daemons, a dead
containerised node — those are observations, not intentions.

Records 0001-0026 are immutable. A record says what was decided when it
was decided, and no record is edited for a name.

Also repoints ADR 0022's link at the renamed skill — a path fix, which the
immutability rule permits, not a change of meaning.
2026-08-23 21:26:09 +02:00
jschoubben daf3e17c32 self-hosting, provisioning and delivery efforts, and the dotfiles origin
The identity provider is settled as not-substrate: the mesh does not
require one, tier 2 authenticates natively, and it is a hosted service
like any other. Four substrate services, not five. The tier test's second
step gains the verb that matters — can the control plane START without it,
not function fully without it.

That verb answers the forge and the registries. They are not substrate and
they are not duplicated: the control plane starts and manages nodes
without a forge, it just cannot change itself. One gitea module, tier 4,
and the mesh's own instance is distinguished by what it is bound to rather
than by being a different module — the same answer as postgres, from the
same test. It also buys a property worth having: if the forge dies the
mesh keeps running.

Delivery needing them is not an upward dependency, resolved the way the
constitution already says to: tier 2 declares requirements, tier 4
provides implementations, the binding is data. The mechanism is
provisioning, and the new idea is that the control plane is itself a
consumer.

Self-hosting therefore becomes a state the mesh REACHES, not a
precondition. A first node comes up from pinned external artifacts and
re-binds to internal providers once they exist. Today's mesh assumes the
second state from the first moment, which is why the first-node path needs
a script that papers over an impossibility and is the least-exercised code
in the system. Made explicit, the transition is also reversible.

Research 007 and 008 opened for the two areas flagged as important and
complex, scoped from the weaknesses the as-is layer already documents
rather than started blank.

And the origin: this began as a dotfiles repository. The first two days
adopt dotfiles, add per-node overrides, and introduce service symlinking
with an ignore file. The flat one-directory-per-tool catalogue, linking
over copying, adoption of already-configured machines, per-node overrides
and the desktop modules are all inherited rather than chosen for a mesh.
That is the single most useful fact for anyone changing the catalogue, it
strengthens ADR 0018 — the case for links was never made for a mesh — and
it explains research 005's silent fifty: dotfiles-era entries for one tool
never shared a domain because they never had one.
2026-08-23 20:55:14 +02:00
jschoubben 3f6d939930 Every decision is a record; the ledger is gone
papa-hq has no ledger. Its root is AGENTS.md, CLAUDE.md, README.md, every
decision is a numbered record, and its graduation playbook has no path for
an unrecorded decision. hal-hq now matches.

The ledger's 41 entries classified as: 10 restating a record, 11 restating
design docs, 15 describing how this repository works with the reasoning
sitting in a README rather than anywhere citable, 3 small rules with no
home, 2 superseded stubs. Mostly a copy — and a hand-maintained index, the
exact pattern ADR 0022 had just rejected for the decision index on the
grounds it drifted after one addition. Keeping one copy of that while
removing another is not a position. It also collided by name with
02-DECISIONS/ in any directory listing.

Nothing was dropped. Records 0019-0025 give the repository decisions the
reasoning they never had: HQ is its own repository and is public, design
has two layers, work moves through playbooks, status lives in frontmatter,
issues have a front door, the numbering is the flow, HQ is the source of
the constitution. 0026 records the ledger's own removal.

The three orphan rules went to how-we-build, where a rule is enforced and
keeps the incident that earned it — the package rule was genuinely
unwritten anywhere. Two lab decisions stated only in the ledger went into
the lab design. "Deliberately not decided" went to the research effort and
design document each question actually belongs to.

The chronological view the ledger provided is now generated from record
frontmatter, which is what it was for.

The cost, stated in 0026 rather than glossed: a record is more work than a
table row, so the risk is a small decision going unrecorded because nobody
wanted to write a document. how-we-build takes rules cheaply, which is the
mitigation, not a solution.
2026-08-23 18:17:59 +02:00
jschoubben c0b35652d0 The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
2026-08-23 18:05:11 +02:00