Commit Graph
71 Commits
Author SHA1 Message Date
jschoubben 902739acb6 Research 011 — the module graph
The proposal to split modules into provisioning services and applications was
worked through and abandoned, for a reason worth keeping: it cannot be filed
consistently. A git forge is consumed as a service and operated through a web
interface; an analytics service grants tracking identity and is a dashboard.

The operator's correction is the sharper form — what runs on the machine is a
supervised container, not something a user started. That is a fact about HOW a
thing runs, not about what kind of thing it is. So it is a facet, and 0002
survives: everything is a module.

What the catalogue is missing is not a taxonomy but a graph. Grouping asserts
relationships; a graph records them. Five declarations, of which two exist:
requires/provides a resource (yes), requires/excludes another module (no),
requires a node capability (no). Plus interface modules that carry no
implementation, with adapters providing them.

Recorded because it matters: this is a package manager's model, and pacman
already has all of it — depends, conflicts, and provides as virtual packages,
which is exactly the interface/adapter idea. Arriving there independently is
evidence for the shape. It is also a warning about what not to reimplement.

Working position on capabilities, to be tested: intrinsic ones (hardware,
architecture, network position) are detected and never installed, and a module
requiring one it lacks is impossible rather than unresolved. Provided ones (a
display server, a container runtime) are not a separate kind of thing — they are
modules that provide a capability, so "may the mesh install a capability" is not
policy, it is dependency resolution. Issue 007 then bears directly: an installed
package is not a capability.

Also captured: the operator's assessment that the machinery around a module —
scheduled tasks, hooks, migrations, config and env — is worth keeping, seeds are
not, and the integration is wrong enough to need a major refactor. Research 005
found supporting evidence from another direction, that the densest apparent
coupling in the catalogue is manifest boilerplate churn.

The first open question is the one that decides whether this is progress: what
does the graph DELETE? If modules gain declarations and lose nothing, it is
motion.
2026-08-25 23:19:57 +02:00
jschoubben 72b22830f3 ADRs 0036, 0037, 0038 — what a node is, what the host does, how one joins
0036 (accepted): a node is a managed machine, and disconnection is a situation.
The open question posed a class distinction — full nodes and lesser presences.
There is none. Reachability is state, not kind, which promotes the host's local
store from a component to a requirement: it is what makes disconnection ordinary
rather than exceptional. The reduced contract the question reached for is real
but it is capability, and that belongs in the profile.

0037 (accepted): the host applies, it does not decide. Measured rather than
argued — the absorption is smaller than the machinery that already applies
state, and eight of ten adapters carry no dependency to move. The two that do
open a Postgres connection to the control plane, which inside tier 0 is the one
thing the tier rule exists to forbid. So each concern splits: deciding needs
every other node and stays in tier 2; applying needs root and locality and goes
to tier 0. The host carries ONE concern, of which the six are instances.

0038 (proposed): a node joins by linking first. The operator's two-modes
proposal, adopted as intent and corrected as structure. Two modes is two code
paths where the first runs once per mesh and rots — and the mesh already has
that fault in its worst form, as three hand-run shell scripts. Instead: one
behaviour, two sources of declaration. The first node is not a different kind of
node, it is a node whose mesh is not up yet, and its specialness is temporary
and self-erasing.

0038 also shrinks the migration 0037 called expensive: a joining node never
needs mesh-wide state, because the hard part of the overlay is only needed to
compute the WHOLE mesh. It needs one peer. The rest arrives.

Left open and said so: what may be pushed over the link and how a joining node
proves it is entitled to join, and whether one host can raise the substrate
alone.
2026-08-25 10:34:45 +02:00
jschoubben 42bce02bba 006: answer the host-size question by measuring it
The skeleton's biggest unproven claim was that absorbing six concerns makes a
binary whose whole argument is having no dependencies carry six of them.

Measured against origin/main, and the question turns out to ask about the wrong
axis. By size the absorption is SMALLER than the machinery that already applies
state on a node — 2755 lines of adapters against 3059 lines of meshware,
env-sync and config-sync. The host is not a new large thing; it already exists,
spread across three core modules.

The real risk is direction, and it is two modules wide rather than six concerns
wide. Eight of ten adapters already receive derived state and only apply it, so
absorbing them moves code that has no dependency to move. Two — wireguard and
traefik — open a Postgres connection to the control plane and compute their own
configuration, which inside tier 0 would be an upward dependency and is exactly
what the tier rule forbids.

And the split has already been happening without being named: dnsmasq-app needs
the same node data as wireguard and does not query for it, because hand-
duplicated state went wrong and someone derived it centrally instead. Eight of
ten adapters are on the far side of that migration.

So the absorption is not a move, it is a split: deciding stays in tier 2,
applying goes to tier 0. The claim survives with its scope corrected — the host
carries ONE concern, apply declared state on this machine, of which the six are
instances.

Stated open rather than glossed: the two unsplit modules are the two hardest,
six concerns is still six vocabularies even at zero dependencies, and what the
host must carry versus find is issue 007 and unresolved.

Question B also recorded as answered by the operator — a node is a managed
machine, and a disconnected node is still a node in a different situation. The
question posed a class distinction; there is none, and what varies is state.
2026-08-25 02:20:32 +02:00
jschoubben 4bf7a35568 Close the record on the lab
Playbook 02 and 04 were followed for the substance — decisions before design,
design before build — and skipped for the bookkeeping. This closes that.

004 graduates. Its one open item was "not yet stood up"; the lab is stood up,
and the substitution the effort turned on is now enforced by the validator
before anything is raised rather than left as a thing to remember. Its
certificate conclusion has a home in 01-end-to-end-testing and is designed but
not built — implementation is a third axis, and an effort graduates on its
conclusions.

One item leaves 004 without a home and is recorded rather than lost: the reverse
proxy does not set caServer, so it defaults to the production endpoint.

The two lab designs read `designed` while running in production of a sort, so
they become `in-progress`.

And the lab gets an as-is document, which it did not have. It records what runs
including the parts nobody would choose again: that `place:` is refused and the
lab therefore raises EMPTY MACHINES, that the drawing shipped with no design
document behind it, that a router is tagged as a machine for a reason found by a
bug, and that the integration suite raises two of five scenarios while both
faults found so far lived in the three it does not.

006 stays active, deliberately. Two of its open questions ARE the tier 0 design
— whether absorbing six concerns makes the host too large, and whether an
unprivileged node earns a place in the inventory. Playbook 04 is explicit that
an open question is a reason to research, not to build around.
2026-08-25 01:55:52 +02:00
jschoubben e88b448145 The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing
1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle
falls from ~90s, unbounded at worst, to ~15s dominated by a boot that
cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write
and did not without it.

The consistency matters as much as the speed: three consecutive snapshots
took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never
finished.

One honest counter-observation recorded: launching onto the fresh
copy-on-write pool was slower, 20.2s against 14.3s, because the image had
to be unpacked into a pool that had never seen it. Paid once per pool, and
dwarfed by what snapshotting saves, but it went the other way.

Doing the measurement produced the answer to how the lab installs on a
clean machine, because both failure modes appeared while doing it.

Installed is not available: the daemon was present with units disabled and
no group. Issue 007.

Available is not adequate, and this is worse: with the storage tooling
absent everything worked and snapshots were seventy-six times slower.
Nothing failed, nothing warned. That is a variant the mesh has not
catalogued — its usual failure is reported success and did nothing; this is
reported success and did it seventy-six times slower, which no error
surface catches because nothing is wrong.

So the lab verifies CAPABILITY, never installation, and refuses to run
degraded rather than warning — a warning about a slow inner loop is read
once and ignored forever. Prerequisites may arrive from a mesh module or
from the lab's own bootstrap, and the second path is required rather than
convenient: a lab installable only by a mesh cannot host the development
of the mesh that installs it.

The lab is the second thing installed by hand, after the node host, and for
the same reason: something has to be first, and pretending otherwise
produces a circularity papered over by a script nobody exercises.
2026-08-24 00:14:57 +02:00
jschoubben 98bcd5cc49 Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than
arguable, so it was measured. One virtual machine on a workstation with
hardware virtualisation and NVMe.

Raising: the launch call returns in 3.4s, the machine is actually usable
after 14.3s. The gap is a design constraint — raise must wait for the
second number, because reporting the first would be transport reported as
effect, which is the mesh's own recurring failure.

Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full
copy; nothing is shared. Restore: 10.4s, usable again after 20.1s.

The second snapshot exceeded two minutes and never completed. That is the
more troubling number: snapshot cost here is not merely high, it is
unpredictable, and a loop with a variable multi-minute step is one nobody
trusts.

Projected to a four-machine scenario, a reset-and-rerun cycle is about a
minute and a half at best and unbounded at worst, before any of the mesh's
own work begins. That is too slow for an inner loop, and ADR 0029's whole
argument — that making the bootstrap path the inner loop turns the
least-exercised code into the most-exercised — holds only while resetting
is cheap.

The cause is not virtual machines. Hardware virtualisation is present and
machines boot in fourteen seconds. It is that the daemon offers exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap
snapshot. The btrfs kernel module is available; btrfs-progs is simply not
installed, which is the entire reason the driver is absent.

The copy-on-write comparison was deliberately NOT run, because running it
would mean installing a package by hand — which the rules forbid and which
would have made the measurement unreproducible. So the honest statement is
that the current configuration is too slow and the likely fix is known but
unverified, rather than that btrfs fixes it.
2026-08-24 00:07:22 +02:00
jschoubben a72fea5342 ADR 0031 and the scenario declaration
The lab provides the underlay; the mesh builds the overlay. This is the
boundary that decides whether the lab is worth having: a scenario that
assigns overlay addresses, elects the hub and writes peer configuration
certifies its own work — if the mesh's peering is broken, that scenario
still comes up green. The most valuable thing the lab can test is exactly
the part pre-building would replace.

So a scenario declares what a hosting provider and a home router would
provide: segments, which machine sits where at which address, what NAT is
between them, which ports are forwarded, which machines are detached. It
declares nothing about overlay addresses, hubs, peering, names or
certificates, all of which become outcomes to observe.

The declaration has four parts — segments, machines, place, snapshot — and
the two scenario classes differ only in place. That is what makes one a
strict subset of the other rather than a fork.

Research 004's most important finding becomes a format constraint rather
than a footnote: the routable segment must use RFC 5737 documentation
space, because the mesh decides public versus private by matching the
address, and a private range there makes the hub test as unreachable while
the mesh silently never forms. A segment without behind: is routable, and a
non-documentation address in it should be refused before anything is
raised — ADR 0008 applied to a configuration file, since the failure it
prevents has no error at all.

Four things left open, including the one that matters most: a lab machine
is always privileged, so the user and edge profiles have no scenario that
exercises them.
2026-08-23 22:33:52 +02:00
jschoubben 09489a298c ADR 0030: the repository structure, and the rule that names them
The tiers were settled and the product was named, but the repositories
themselves existed only in a research sketch. That had already caused two
problems.

ADR 0029 makes the lab phase 0 of the migration and could not say where it
lives, because no record named a repository.

And the sketch contradicted an accepted record: it listed mesh-hq while
ADR 0028 had decided novox/hq and explicitly rejected that name. A design
resting on research is resting on something that can change without a
decision. Corrected in the research too.

The naming rule, which both earlier records implied and neither stated: a
repository belonging to a product carries that product's prefix; a
company-scoped one does not. That is why this repository is hq and the
mesh's are mesh-*.

Seven repositories recorded — host, substrate, control, surfaces, sdk, lab,
and this one. The lab gets its own: it ships to nobody, outlives any single
tier, and drives virtualisation on a workstation, which nothing else does.
Inside the host it would couple development tooling to a shipped
component; inside the control plane the bootstrap scenario would depend on
a tier that does not exist when it is needed.

Tier 4 is deliberately not decided. Whether the catalogue is one
repository, one per domain or one per application stays open from ADR 0015
and is blocked on research 005 — how many repositories hold domains cannot
be answered before knowing what the domains are. mesh-catalog appears in
the sketch and is not decided by this record.

The cost is stated rather than glossed: seven release cadences where there
is one, and cross-repository changes that used to be one commit.
2026-08-23 22:04:04 +02:00
jschoubben b4904fec7e The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a
complete mesh — forge, coordinator, cascade, verify. That is unusable for
building the new mesh, because all four are tier 2 and do not exist yet.

And research 009 had the sequence backwards. It placed the lab at phase B
as verification of tiers already built, but tier 0 is the component that
takes over a machine's packages, services and network. It cannot be
developed against a machine anyone needs. The lab has to exist before the
thing it will test.

ADR 0029 splits scenarios into two classes. The bootstrap scenario is
virtual machines, the host binary and a pinned bundle, with the verdict
coming from what the host reports about the state it reconciled. The full
scenario is the designed one. The first is a strict subset of the second —
same virtualisation, same networking, same lifecycle, stopping before a
control plane exists — so the second is reached by addition rather than
rework.

The consequence worth having: raising a node from nothing stops being the
least-exercised path in the system and becomes the inner development loop.

It also settles the runner's two jobs. Scenario lifecycle is needed
immediately, because something must materialise and reset a mesh before
anything can be written against it. Assertion execution waits for the full
scenario.

Corrects a stale claim in the design while amending it: it argued
scenarios were affordable with system containers and would not be with
virtual machines. ADR 0016 superseded that reasoning and the text had not
followed.

Issue 007: the lab's first requirement is installed and unusable. The
virtualisation package is present and explicitly installed; both units are
disabled, the operator is in no group, and the client reports the server
unreachable. Not issue 001 again — that is an install failing while
reporting success. This is an install succeeding when success was not the
point. A package is files; a capability is a running service and an
identity permitted to reach it, and the module model has no vocabulary for
the second.
2026-08-23 21:57:14 +02:00
jschoubben c0ae8dec96 Remove two disclosures, and record Nox as the answer to 006
Found by a full scan before making the repository public, which is the
moment the public rule stops being aspirational.

A module name identified a specific laptop model — hardware inventory,
which is operational detail about one installation rather than a lesson
that travels. Generalised.

ADR 0028 named a forge username in a repository path, which the public
rule forbids, and the sentence had also gone stale: the repository it
described was subsequently verified empty of anything unique and removed.
Rewritten to state what happened without the username. Removing a
disclosure from a record is the same class as fixing a path — the rule
that permits it outranks the one that forbids editing.

Issue 006 gains its proposed direction: Nox works from within this
repository rather than these documents being synced into the knowledge
base. Better on three counts — no copy, so no drift; no fourth knowledge
system, which was the original objection; always current.

But it changes the promise, and the issue says so. ADR 0019 promised these
documents would surface BESIDE everything else in a symptom search. An
agent that must be asked is reachable, not surfacing, and the two differ
in precisely the case the operational memory exists for — someone
debugging an error with no reason to suspect HQ knows anything about it.
The question narrows to whether a symptom search finds this content
without the searcher already suspecting it.
2026-08-23 21:40:53 +02:00
jschoubben 87f4f29cc6 Novox Mesh, Nox, and HQ becomes company-scoped
ADR 0027 — the product is Novox Mesh, shortened to mesh internally. HAL was
never chosen: it arrived with the dotfiles repository this grew out of, it
is borrowed, and it is borrowed from the canonical untrustworthy machine
intelligence, which is an odd flag for infrastructure trusted with
credentials. Timing is the substance of the decision, not an aside — the
skeleton is not built, so renaming costs a search and replace now and a
migration later.

Nox is an identity of Novox, and specifically the agent of the MESH rather
than of a node. Nodes keep their own identities. Nox addresses them, and a
human mostly talks to Nox — which makes it the concrete form of the
mission's vision: state an intent, and the mesh works out which node holds
the thing. It holds no private channel. The gap this opens is recorded:
ADR 0012 binds every agent to a home node, and a mesh-scoped agent has
none, so the model needs extending.

ADR 0028 — HQ is company-scoped, novox/hq, with the mesh as its first
product. Checked rather than assumed: the company organisation already
holds live projects that the mesh builds and deploys, so they are tenants
rather than peers, and the mesh is the ground they stand on. There is also
company work outside the mesh already, which strengthens the case and means
the eventual split is closer than "some day" — so each document's scope is
fixed now, in a table, making that split mechanical instead of
archaeological. The folders are deliberately not restructured yet.

The skeleton takes the new vocabulary: mesh-host, mesh-substrate,
mesh-control, mesh-surfaces, mesh-catalog. Substrate drops to four services
now that identity is a hosted workload rather than a dependency.

Research 009 opens the migration, with the reframing that lowers its risk:
replace the control plane, do not move the workloads. Their data never
moves, so it is re-declared rather than adopted — which keeps adoption out
of scope, as the lab design requires. Self-hosting is the last phase, or a
failed cutover takes away the means to fix it.
2026-08-23 21:20:35 +02:00
jschoubben daf3e17c32 self-hosting, provisioning and delivery efforts, and the dotfiles origin
The identity provider is settled as not-substrate: the mesh does not
require one, tier 2 authenticates natively, and it is a hosted service
like any other. Four substrate services, not five. The tier test's second
step gains the verb that matters — can the control plane START without it,
not function fully without it.

That verb answers the forge and the registries. They are not substrate and
they are not duplicated: the control plane starts and manages nodes
without a forge, it just cannot change itself. One gitea module, tier 4,
and the mesh's own instance is distinguished by what it is bound to rather
than by being a different module — the same answer as postgres, from the
same test. It also buys a property worth having: if the forge dies the
mesh keeps running.

Delivery needing them is not an upward dependency, resolved the way the
constitution already says to: tier 2 declares requirements, tier 4
provides implementations, the binding is data. The mechanism is
provisioning, and the new idea is that the control plane is itself a
consumer.

Self-hosting therefore becomes a state the mesh REACHES, not a
precondition. A first node comes up from pinned external artifacts and
re-binds to internal providers once they exist. Today's mesh assumes the
second state from the first moment, which is why the first-node path needs
a script that papers over an impossibility and is the least-exercised code
in the system. Made explicit, the transition is also reversible.

Research 007 and 008 opened for the two areas flagged as important and
complex, scoped from the weaknesses the as-is layer already documents
rather than started blank.

And the origin: this began as a dotfiles repository. The first two days
adopt dotfiles, add per-node overrides, and introduce service symlinking
with an ignore file. The flat one-directory-per-tool catalogue, linking
over copying, adoption of already-configured machines, per-node overrides
and the desktop modules are all inherited rather than chosen for a mesh.
That is the single most useful fact for anyone changing the catalogue, it
strengthens ADR 0018 — the case for links was never made for a mesh — and
it explains research 005's silent fifty: dotfiles-era entries for one tool
never shared a domain because they never had one.
2026-08-23 20:55:14 +02:00
jschoubben 7a20358113 research 006: the code skeleton, and where postgres lands
A tier test as a decision procedure — five ordered questions, first match
wins — so placement is answerable rather than argued.

Postgres was the test case and the naive answer is wrong. Not twice, once:
the control plane cannot exist without a relational store, so it is tier 1
and lives in hal-substrate/store/postgres. What differs between the mesh's
own database and a project's is not the module but how that instance is
brought up — pinned bundle applied by the host, versus the ordinary
delivery and provisioning path. Tier is a property of the module; the
bundle is a property of the mesh's own instance. The naive answer would
also have made substrate reach up into the catalogue, which the dependency
rule forbids.

Working the test across the catalogue surfaces a third fate that neither
of research 005's options covers, and it is the most common one: absorbed
into the host, ceasing to be a module at all. That explains 005's one
positive measurement rather than confirming it — the reachability cluster
is not four modules that should be one domain module, it is four facets of
one thing the host should own, expressed as modules because a module was
the only unit available. Under this skeleton the overlay and firewall
modules stop existing. It also partly answers the silent fifty: several
are host concerns, so silence was the right signal and grouping was the
wrong inference.

Flags rather than settles: the identity provider is a genuine boundary
case (four substrate services or five), and absorbing six concerns into a
binary whose argument is that it has no dependencies is the skeleton's
biggest unproven claim.
2026-08-23 20:40:32 +02:00
jschoubben 00b8398d07 skeleton: hal-agent -> hal-host, and the two gaps the questions found
Agent is a first-class concept here — a participant, some of whom are
human, holding identity and memory. Using it for the tier-0 node binary
put both meanings in one document: hal-agent at tier 0, agents at tier 2.
That is the anatomy-naming failure again, an evocative domain word
pointing at infrastructure, and how-we-build 4 exists to catch it.

ADR 0015's own title supplies the fix: the mesh brokers, nodes host,
agents think. So the control plane brokers (hal-mesh), the tier-0 binary
hosts (hal-host), the participant thinks (agents, untouched). hal-node was
rejected — Node is the inventory aggregate, the binary is what runs on it.
Recorded in the document as a near-miss rather than quietly corrected.

Tiers now explained before the tree, as a boot narrative, with the point
they were carrying made explicit: dependencies point only downward, that
is the whole bootstrap answer, and the current mesh violates it — the
database is a module, modules come from the pipeline, the pipeline needs
the database.

Connectivity gains the part that was missing. Naming a context says who
decides, not who runs it: overlay membership is tier 0 in the host, policy
is tier 2, machinery is tier 4 modules. And the host's link to the control
plane deliberately does not run over the overlay, or the overlay would
have to exist before a node could be told how to join it.

'tier' is this document's coinage and did not land on first reading; that
is now an open question rather than settled vocabulary.
2026-08-23 20:32:20 +02:00
jschoubben b4365d8aa4 research 006: the mesh designed from nothing
A skeleton laid out against the stated requirements rather than derived
from the current shape: four tiers, repositories at the root, and a
dependency rule that only points downward.

Four moves the current shape does not have.

The substrate is applied by the agent from a pinned bundle, not delivered
by the pipeline. That is the bootstrap circularity removed rather than
worked around — the first node is the ordinary path with no control plane
on the other end, which also makes it the cheapest lab scenario instead of
the one nobody exercises.

One agent binary with a detected capability profile — managed, user, edge.
A phone becomes a capability question rather than a platform question, so
it needs no second implementation. Modules declare which profiles they can
land on, and an impossible assignment fails at declaration.

Connectivity becomes a context. ADR 0015 names nine and none owns the
overlay, resolver, firewall or ingress, while research 005 measured
reachability as the only cluster in the catalogue that genuinely changes
together under one intent. Gap and evidence point the same way. That is an
addition to an accepted record, so it needs its own record and is not
written here.

Feature splits into artifact (built once per version) and part (selected
per node). The conflation of those two cardinalities under one word is
what makes the delivery pipeline hard to reason about.

Also makes explicit in how-we-build that the main-branch rule covers this
repository too. The rule already said 'without exception'; nothing was
amended, so nothing is recorded.
2026-08-23 19:25:29 +02:00
jschoubben 143f8ab2f1 research 005: which modules actually change together
The domain-grouping premise is testable, so it was tested before drawing a
list. Co-change across the full history of the module catalogue, current
modules only, platform namespace excluded.

Nine commits in ten touch exactly one module, and 50 of 89 modules have
never been edited alongside anything. Two clusters exist above that floor.

Reachability holds up: proxy, resolver, firewall and VPN genuinely move
together under one intent, three times in recent history. That is the shape
ADR 0017 describes and the only place the measurement finds it.

The provider cluster does not, and this is the finding worth having. Every
multi-provider commit is a cross-cutting manifest change applied N times —
feature detection, hook conventions, volume binds, network scoping. None is
a change to what a database is. Merging them would not have prevented one
of those commits, and the history already shows the fix that worked:
verify by shape in the SDK rather than copying a script into every module.
Move the concern into the machinery, do not merge the modules carrying it.

ADR 0017 keeps its principle and gains a pointer to this narrowing.
2026-08-23 18:48:28 +02:00
jschoubben 3f6d939930 Every decision is a record; the ledger is gone
papa-hq has no ledger. Its root is AGENTS.md, CLAUDE.md, README.md, every
decision is a numbered record, and its graduation playbook has no path for
an unrecorded decision. hal-hq now matches.

The ledger's 41 entries classified as: 10 restating a record, 11 restating
design docs, 15 describing how this repository works with the reasoning
sitting in a README rather than anywhere citable, 3 small rules with no
home, 2 superseded stubs. Mostly a copy — and a hand-maintained index, the
exact pattern ADR 0022 had just rejected for the decision index on the
grounds it drifted after one addition. Keeping one copy of that while
removing another is not a position. It also collided by name with
02-DECISIONS/ in any directory listing.

Nothing was dropped. Records 0019-0025 give the repository decisions the
reasoning they never had: HQ is its own repository and is public, design
has two layers, work moves through playbooks, status lives in frontmatter,
issues have a front door, the numbering is the flow, HQ is the source of
the constitution. 0026 records the ledger's own removal.

The three orphan rules went to how-we-build, where a rule is enforced and
keeps the incident that earned it — the package rule was genuinely
unwritten anywhere. Two lab decisions stated only in the ledger went into
the lab design. "Deliberately not decided" went to the research effort and
design document each question actually belongs to.

The chronological view the ledger provided is now generated from record
frontmatter, which is what it was for.

The cost, stated in 0026 rather than glossed: a record is more work than a
table row, so the risk is a small decision going unrecorded because nobody
wanted to write a document. how-we-build takes rules cheaply, which is the
mitigation, not a solution.
2026-08-23 18:17:59 +02:00
jschoubben c0b35652d0 The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
2026-08-23 18:05:11 +02:00
jschoubben f05e4a0dce Follow papa-hq's research convention; the mesh links nothing
Research efforts move from status.md to 00-overview.md with active /
graduated / abandoned, matching papa-hq so the two repositories read the
same way. Playbooks, skills, README and the ledger follow.

Reverses yesterday's withdrawal of the symlink note in GENESIS. The note
was right and the withdrawal was wrong: the intent is that the mesh
creates no symlinks at all, so a founding document listing "symlinks, not
copies" as a design principle does point the opposite way from where this
is going, and that is a contradiction rather than a stale detail.

ADR 0018 records the position, proposed. ADR 0011 stays as it is — it is
the historical decision and the incident behind it is why anyone believes
either record — and is superseded in intent, not edited. Its one
editorial line, which called the wider reading false, is corrected to
state what is actually true: centralising who may link narrowed the
incident class without closing it, because a link the installer makes
resolves exactly like one made by hand.

The argument that kept linking was staleness. ADR 0004 removed it: every
managed file is already derived and reconciled, so a copy is the natural
form and a pointer into source is the shape the mesh's own model forbids
everywhere else. What is not settled, and is marked open, is how
staleness gets detected — which is the decision that makes or breaks it.
2026-08-23 09:29:09 +02:00
jschoubben 702efca6bb Base layer: the mesh as it is, under the mesh as it should be
HQ held only the to-be. Every reader had to already know the system the
decisions were about, and an as-is claim had nowhere to live except inside
an intention.

Adds 02-DESIGN/00-as-is — eleven documents written from the implementation
and the operational record, not from intent, including the parts nobody
would choose again. The two existing designs move under 01-to-be. Layers
are declared in frontmatter and never mix: a design that ships does not
move, its as-is counterpart is written, and both stand.

Back-fills adr/0001-0014 for decisions taken in implementation and never
recorded — the broker, the module abstraction, the mesh database, managed
files, provisioning, migrations, the workspace removal, failing loudly,
the constitution, application placement, linking, the employee model, the
artifact, the three silos. Each marked reconstructed, dated from the
history, and citing the evidence it was recovered from. The two existing
records renumber to 0015 and 0016 so the ledger runs oldest first;
0017 extends 0015 to modules outside the core, principle only — the
domain list is deliberately not invented here.

how-we-build.md becomes the source of the mesh constitution, with a sync
playbook, so the enforced copy stops being the only one that is true.

Process becomes explicit: five playbooks, eight thin skills that defer to
them, a repository map, and AGENTS.md with CLAUDE.md as its include.

The five Observations become 04-ISSUES 001-005 where they can be owned and
closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge
base. That claim is what decision 27 rests on, it was never checked, and
the README now says so instead of repeating it.

Also corrects the ADR index into something generated, the "02-DESIGN is
empty" claim, the VISION.md pointer that did not survive the repo split,
and a note asserting the symlink rule was contradicted — it was a
misreading; the rule forbids hand-made links, the installer links by design.
2026-08-23 03:08:26 +02:00
jschoubben cf9357e8e9 HQ — the mesh's own documentation
What the mesh is, what it is becoming, and why. Implementation lives in the
code repositories; the reasoning lives here.

  00-GENESIS   mission, engineering context, effect, and the rules that hold
  01-RESEARCH  investigations, before they harden into design
  02-DESIGN    the authoritative specification
  adr          numbered decisions — what was chosen, and what was rejected
  DECISIONS.md the ledger: every decision, in the order it was taken

Written for a reader who is not its author and has no access to the mesh it
describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918;
nodes are named by role.

Single initial commit by intent. The prior history came from a private
repository and carried operational detail — a routable address identified as a
VPN hub, real domain names, a hosting provider — which sanitising a tip commit
would not have removed from the log.
2026-08-22 22:01:32 +02:00