Commit Graph
861 Commits
Author SHA1 Message Date
jschoubben e10084fed6 Merge pull request 'Issue 087: a machine states its host only when the mesh sends it something' (#205) from issue/087-a-machine-says-its-host-only-when-asked into main 2026-09-30 07:09:17 +00:00
jschoubben 941b920bd7 Issue 087: a machine states its host only when the mesh sends it something
Measured live. All four machines run the identical host binary — same
digest, installed within eighteen seconds — and at first only one reported
a version, which read as a difference where there was none.

A machine publishes a report after an apply. The five-minute reconcile
publishes nothing, because it is the machine keeping itself as declared
rather than answering anything. So a current, idle machine never says, and
the mesh cannot tell that from a machine running something ancient.
Confirmed by pushing: not reported, then 04a27ca.

Enough for the purpose, not enough for the claim. For deciding whether a
new declaration field is safe it is sufficient — pushing is what the mesh
is about to do, and the answer arrives with the act. For knowing what the
mesh runs it is not, and "not reported" is worded as "nobody has asked
recently" for that reason.

Making a heartbeat carry it would close the gap and would change what a
heartbeat is — a bare word that the node is there, deliberately carrying
nothing else. Left alone rather than widened in passing.
2026-09-30 09:09:10 +02:00
jschoubben ecbd2ce2e6 Merge pull request 'Group 1: 145's report states its scope, and 107 waits for delivery' (#204) from issue/145-and-107-what-group-one-leaves into main 2026-09-30 07:00:32 +00:00
jschoubben b7f7b97d8a Group 1: 145's report states its scope, and 107 waits for delivery
145, partly resolved. The sentence that was true for eleven hours of a
mesh in which no module could reach another now says what it is not a
claim about: that is the mesh and the machines agreeing, and nothing here
dials a provision. It checks nothing and does not pretend to — ADR 0146
decides the check and is deliberately not built. What changed is that the
report no longer implies otherwise. Stays open for that reason.

Carried forward: 0146's check needs an internal name fetched over TLS with
the certificate verified, and until today no machine trusted the mesh's
authority. Three of four do now, so whoever builds it does not have to
solve that first.

107, diagnosed and deliberately not built. The premise is confirmed in the
host's own words — unknown fields are refused because "a field the host
does not know is a thing the control plane believes it asked for" — so the
fix is a flag day, not an addition. 087 now makes the cost measurable, and
the measurement is why it waits: one machine of four runs an older host,
it is parked, and nothing delivers a host at all (142). Shipping the field
means hand-placing binaries and unparking a machine, and one missed in
that sequence is unreachable, not degraded. The fault it prevents has
never been observed.

142 gains the note that it is 107's gate, and that it is what makes a
declaration field cost a rollout instead of an expedition.

A judgement about order, not a refusal, and cheap to overrule.
2026-09-30 09:00:18 +02:00
jschoubben c2cf0d72d6 Merge pull request 'Issue 087 is resolved: the mesh knows which host runs a machine' (#203) from issue/087-the-mesh-knows-which-host-runs-a-machine into main 2026-09-30 06:55:05 +00:00
jschoubben 0a5006b366 Issue 087 is resolved: the mesh knows which host runs a machine
The machine has reported its host version since ADR 0141, whose own
comment says why it must: without it nothing can say a machine is behind.
The controller's copy of the report did not have the field, so it
unmarshalled into nothing and was thrown away on arrival. Two structs
describe one message and only the sending side had it.

node show names it per machine, "not reported" where the mesh has not been
told. status names every machine running an older host than another does,
and which is newest.

Disagreement rather than staleness, deliberately: nothing delivers a host
version yet, so the mesh holds no canonical current one and "behind" has
no fixed point. What it can say is that the oldest host in the mesh is
what the mesh may send.

Two refusals to guess: a machine that reported nothing is not called
behind, and versions compare as strings — right for the timestamps this
mesh uses, wrong for a scheme where 10 sorts before 9, said at the place
that would have to learn.

107 gains the note that this is what makes its new field safe to consider,
and that one machine of four is behind today, so it is not free yet.
2026-09-30 08:54:52 +02:00
jschoubben 3ea9601904 Merge pull request 'Issue 125 is resolved: a hold is a line in the report, and it breaks "all well"' (#202) from issue/125-a-hold-is-a-line-in-the-report into main 2026-09-30 06:47:57 +00:00
jschoubben f7a37ee4f3 Issue 125 is resolved: a hold is a line in the report, and it breaks "all well"
Two of the four surfaces the report named already carried it — the host
has reported Held since ADR 0100, and node show reads the machine's own
list with an `as of` beside it. Recorded as checked rather than assumed.

Two did not. The apply line counted what it applied and said nothing
about the difference; status read the mesh's take-time listing, so a
module assigned after it showed nothing at all.

Both now say it, and the part that carries the weight: a hold suppresses
"all doing what they were told, all heard from, running what the mesh
would send them". That sentence was true for the whole outage, and acting
on it is what stopped the predecessor's proxy. Being adopted still does
not suppress it — a mode somebody chose is not a half-finished action.

Status does not call a hold a fault, deliberately. It is correct
behaviour, and a reader trained to see red for something the mesh did
right stops reading.
2026-09-30 08:47:36 +02:00
jschoubben 601d004fcb Merge pull request 'Issue 129: three machines of four trust the mesh, not one' (#201) from issue/129-three-machines-not-one into main 2026-09-30 00:30:30 +00:00
jschoubben d009c3efef Issue 129: three machines of four trust the mesh, not one
Extended after the first was proven. Every converged machine now holds
the anchor and verifies an internal name with a plain client; before,
the two unassigned ones answered 'unable to get local issuer
certificate' and held no entry for the mesh.

ace is excluded on purpose: it is adopted, so a module assigned there is
held rather than run, which is right and is not trust.

Both the resolution and 0147's insight said one machine of four, which
was true for about twenty minutes.
2026-09-30 02:30:23 +02:00
jschoubben 5036b927b9 Merge pull request 'Issue 129 is resolved: a workstation trusts the mesh, and stops when told to' (#200) from issue/129-a-machine-trusts-the-mesh into main 2026-09-30 00:18:51 +00:00
jschoubben 1f72e82b84 Issue 129 is resolved: a workstation trusts the mesh, and stops when told to
Registered ca-trust from the catalogue — it was merged and had never been
registered, which is why "assign it to one machine" had no module to name
— assigned it to the workstation, and verified.

Verified in the form ADR 0147 prescribes, against the authority's own API
so the handshake needs nothing else in the mesh to be right: 200, issuer
Mesh Internal CA, Verify return code 0. Four routed internal names verify
too, and `git ls-remote https://…` works, which is the consequence the
report named.

Removal exercised for the first time. Unassign and push removes the
anchor, empties the trust store of the mesh's authority, and returns the
plain client to the original error; assigning again restores it. That is
the half 0147 claimed and nothing had shown.

One thing it found that is not in the module: removal works only because
the host removes the service before the script. Stopping the unit is what
deletes the certificate and refreshes the bundles, and it needs the script
to still exist. The symmetry rests on an ordering nothing states.

0147's "written, and not yet run" now carries a progressive insight
saying it has run, where, and that it ran on one machine of four.
2026-09-30 02:18:21 +02:00
jschoubben 76fbe323ea Merge pull request 'Issue 118 is resolved: it was issue 135, and umami is healthy' (#199) from issue/118-is-135-and-is-resolved into main 2026-09-29 23:54:39 +00:00
jschoubben 295bf2f3ab Issue 118 is resolved: it was issue 135, and umami is healthy
Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.

It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.

118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.

Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
2026-09-30 01:54:19 +02:00
jschoubben e1b74810a8 Merge pull request 'Issue 129 is live and reproduced, and needs three steps rather than one' (#198) from issue/129-and-what-reproducing-it-found into main 2026-09-29 23:30:48 +00:00
jschoubben e41eed0852 Issue 129 is live and reproduced, and needs three steps rather than one
The certificate is genuine, from Mesh Internal CA, and nothing on the
workstation trusts it — verbatim the error the report gives. The public
name on the same proxy verifies cleanly, which puts the fault exactly
where the report puts it.

What is in the way is not an assignment. `ca-trust` is merged in the
catalogue and has never been registered with the mesh — 39 of 76
manifests are — so there is no module to assign. It dry-runs clean and
needs no artifact built.

Two findings from reproducing it, both their own issues:

157 — every routed name is published with an `.internal` alias that
nothing serves. The hosts file says keycloak.novox.be.internal; the proxy
serves keycloak.novox.internal and refuses the other by name. The first
three names I tried came from the hosts file and failed with a TLS alert
rather than a verification error, which pointed at a regression that had
not happened.

158 — the proxy re-logs all 52 routes every two seconds, 31 times a
minute. The one line that explained 157 sat between two of them.

Also recorded, because it was nearly filed as a defect and is not one:
step-ca publishes roots as /roots.pem, which is PEM, so ca-trust's fetch
and its refuse-a-non-certificate guard are both right. Its other endpoint
/roots returns JSON that contains the text the guard greps for, so the
guard is sound only because of which path is published.
2026-09-30 01:30:24 +02:00
jschoubben ba30286896 Merge pull request 'The pointers back from what yesterday's records changed, which I missed twice' (#197) from decision/the-pointers-back-from-what-these-narrow into main 2026-09-29 22:47:45 +00:00
jschoubben 53b94c51bb The pointers back from what yesterday's records changed, which I missed twice
Three records were left describing a mechanism a new record had moved,
and a reader arrives at them by following a citation: 0066 still said a
routed name is written into every container after 0148 replaced that with
resolution; 0016 still read as though the lab were the test bed after
0149; and issues 109 and 135 said nothing about 0148 ending the copying
that 135's own fix made comparable. Each was a citation leading to the
wrong answer in a record that was not wrong about anything it decided.

This is the second time in one session. The playbook rule I added last
round did not stop it, so the convention is now written where the record
conventions live, with the shape to use and three worked examples — and
with the honest note that it is NOT machine-checked and cannot be from
`extends:` alone: 102 records extend another, 87 have no back-reference,
and that is correct, because extending usually means building on a
context. Making it mechanical means a record declaring the relationship
in frontmatter, which is a schema change and is not mine to decide.

Also: designs 18 and 20 claimed `updated:` dates from before I edited
them, and 117's `fixed-by` gained the commit beside the record.
2026-09-30 00:47:19 +02:00
jschoubben 83791f0921 Merge pull request 'The four open design questions, answered: ADRs 0148, 0149, 0150, and 0114 accepted' (#196) from decision/0148-the-meshs-names-are-resolved-not-copied into main 2026-09-29 22:39:38 +00:00
jschoubben bf4d4e2e7b The three parked questions are answered: 0114 accepted, 0068 superseded, 117 settled
**0114 accepted.** The separation it draws — a consumer's resource and a
credential that reaches it are different things — is a data-loss rule,
and a record naming one should not sit unresolved while the code that
could hit it is written. Not built, and accepting it schedules nothing;
accepted-and-not-built is where 0141 and 0142 already are.

**0068 superseded by 0149, the live mesh is the test bed.** Never built,
and contradicted by practice that was written down nowhere but a handoff
note. The faults that cost the most are faults of a mesh that already
exists — bound consumers, containers made against an older roster, an
adopted machine — and a bed is by construction a mesh that does not.
Issue 156 settles it: that change was verified honestly, on the only path
where it cannot fail. The lab is not retired; raising a mesh from bare is
now its whole job.

**117 answered by 0150.** A module's own code runs as supervised
processes under the module's one account. 0047's argument was for a
runtime per module, not for a container — a unit satisfies it and runs as
an account. The invariant is the account, not the process count: its
worry was a second identity to scope and seal, and processes sharing one
account create none. Designs 18 and 20 now cite it and 0047 carries a
dated pointer, which closes the third disagreement — that neither design
knew the record existed.
2026-09-30 00:39:09 +02:00
jschoubben ec42ee0846 ADR 0148: the mesh's names are resolved, not copied into every container
Answers issue 151. Copying the roster into each container made the roster
part of each container's identity, so one name moving replaced every
container in the mesh — and it never stopped the staleness it was for,
since a copy taken at creation is stale the moment the roster moves (109,
135).

A container resolves through its machine's resolver instead, and nothing
is copied. Staleness stops being possible rather than detected, and a
name's blast radius becomes nothing.

Scoping each container to the names it binds was the close call and is
rejected: it contradicts anything-calls-anything, and leaves the roster
in the digest so the churn returns for a widely-bound name.

Gated on issue 110 — a container on the runtime's default network has no
DNS at all today. Removing the copy first reintroduces 109 and 135
silently on a live mesh. 151 stays open until the code lands; design 08's
file-not-resolver passage is narrowed to the machine's own roster.
2026-09-30 00:34:30 +02:00
jschoubben 5a3dee9e9e Merge pull request 'The records pointed at branches that no longer exist, and two fixes had no sequel' (#195) from issue/pointers-that-resolve-to-nothing into main 2026-09-29 22:29:02 +00:00
jschoubben 47909c6b71 The records pointed at branches that no longer exist, and two fixes had no sequel
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.

Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.

ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
2026-09-30 00:28:35 +02:00
jschoubben 7e5edab8da Merge pull request 'Issue 156: the wider grant has landed, and the notes the fix printed' (#194) from issue/156-the-grant-has-since-landed into main 2026-09-29 22:19:35 +00:00
jschoubben 3649f82204 Issue 156: the wider grant has landed, and the notes it printed
Measured on the broker's own config after the fixed controller composed
it: every node is now allowed _DELIVER.<node>.> as well as the bare
subject. That grant was missing when this was diagnosed, which is why
re-making the consumers would have silenced every machine.

Also records the five consumers it kept and reported, including the build
machine's worker, which is bound the same way and was not anticipated.
2026-09-30 00:18:57 +02:00
jschoubben b400ce2c54 Merge pull request 'Issue 156: moving a consumer's delivery subject stops a running mesh' (#193) from issue/156-a-consumer-that-works-is-not-replaced into main 2026-09-29 21:57:00 +00:00
jschoubben 995c8cb266 Issue 156: moving a consumer's delivery subject stops a running mesh
146's change is correct on a mesh being raised and fatal on one that is
running: the server will not move a push consumer's delivery subject
while a subscriber is bound, and a node is bound to its declaration
consumer the whole time it is up. Reproduced before fixing; an earlier
version of the test unsubscribed first and passed against the broken
code.

The consumer that works is kept and the assertion says so. Not re-made:
the nodes are granted the bare subject only, so re-making would have
silenced every machine.
2026-09-29 23:56:31 +02:00
jschoubben f89aef992d Merge pull request 'Two records shared a number, twice, and every check passed' (#192) from issue/two-records-share-a-number-and-nothing-says-so into main 2026-09-29 21:39:09 +00:00
jschoubben b967ef7be3 Two records shared a number, twice, and every check passed
Two machines filing issues in the same hour both read main correctly and
both took "the next free number". main lags every open pull request —
seven that evening — so they collided twice. The second collision
reached main with records, cycle and index all reporting success.

cycle.py now refuses a tree where two issue folders share a leading
number, and names both. Proven by adding a duplicate and watching it
fail. The colliding records become 153 and 154, renumbered in the branch
that lands last, because renumbering a branch whose author is still
pushing only moves the race.

The check catches a collision; it does not prevent one. Taking a number
still means reading the open pull requests as well as main — issue 155
says so.
2026-09-29 23:38:48 +02:00
jschoubben e417906241 Merge pull request 'Issue 150: a machine's own network is not a reach' (#190) from issue/150-a-machines-own-network-is-not-a-reach into main 2026-09-29 21:37:41 +00:00
jschoubben b1bf895688 Merge pull request 'Issue 149: an adopted machine's data cannot be placed where it is' (#189) from issue/149-adopted-data-cannot-be-placed-where-it-is into main 2026-09-29 21:37:34 +00:00
jschoubben 0d64677c70 Merge pull request 'Issue 152: a node the mesh could not read withdrew its names from every machine' (#191) from fix/152-a-lookup-failure-is-not-an-absence into main 2026-09-29 21:37:16 +00:00
jschoubben 04205c1dc8 Merge pull request 'Issues 147 and 148: a route before its module is taken; a new name recreates every container' (#188) from issue/147-148-found-migrating-ace into main 2026-09-29 21:37:09 +00:00
jschoubben 9dc49cd831 Merge pull request 'Grooming: five issues were fixed and never closed, and one is not' (#187) from grooming/stale-issues into main 2026-09-29 21:37:04 +00:00
jschoubben 66b413a076 Merge pull request 'Issue 140 is resolved, and was resolved before it was read again' (#186) from issue/140-resolved into main 2026-09-29 21:36:57 +00:00
jschoubben fcba05fed9 Merge pull request 'Issue 146: a first node now enrols, and is enrolled twice' (#185) from issue/146-diagnosis into main 2026-09-29 21:36:42 +00:00
jschoubben 18f37c25b2 Issue 152: the loop is metastable, and it cleared at 23:11
It ran 22:46-23:11, five full replacements, and stopped when a pass
happened to read the store during a window it was up. The record said
nothing about it stops; that was wrong. Exiting by luck is the finding,
not a mitigation.
2026-09-29 23:27:51 +02:00
jschoubben 3c2b4fc6b6 Issue 152 is fixed: a lookup failure is no longer an absence
The three gatherers pass over a node whose set does not compose, and
raise anything else. 151 stays open: this removes the false reasons a
roster changes, not the fact that a real change still replaces every
container.
2026-09-29 23:26:20 +02:00
jschoubben 72eaf52867 Issue 152: a node whose plan will not compose withdraws its names from every machine
ace numbered its two records 147 and 148, which this repository already
uses; they become 150 and 151, as 127 became 149.

The new record is why the control node could not stop applying: the
roster alternates between two values because routeNamesInTheMesh
swallows a per-node plan failure, and the roster is part of every
container's identity. The loop closes through the control plane's own
store, which each pass replaces.
2026-09-29 23:13:54 +02:00
jschoubben 741625e725 Merge remote-tracking branch 'origin/issue/147-148-found-migrating-ace' into issue/the-roster-flicker-recreates-every-container 2026-09-29 23:12:57 +02:00
jschoubben c3730b9a23 Issue 150: a machine's own network is not a reach
ADR 0138's reach is internal, public or both. ace's LAN devices (an IoT
switch on mosquitto, unifi access points, plex clients) are none of them;
the only reach that admits them is public, which states the wrong thing and
opens the port to the internet on any machine with a public address.
2026-09-29 23:06:08 +02:00
jschoubben 3a59099c81 Issue 149: an adopted machine's data cannot be placed where it is
ADR 0112 decided that an assignment places a directory and says where an
access's data is; the control plane implements neither. ace's media stack
(a 40 TB ZFS library, plex state on a second disk) cannot be expressed
without writing ace's paths into manifests.
2026-09-29 23:03:11 +02:00
jschoubben 3bd6f34de3 Issues 147 and 148: a route before its module is taken; a new name recreates every container
Both found migrating the first module on ace (searxng), adopted with route-adapter:
147 — assign is not a no-op for a routed module: its route reaches the predecessor's
proxy while its container is held, and the name serves a dead backend until take.
148 — every push that moved the mesh's names replaced every container on novox, twice,
the control plane's own store included.
2026-09-29 23:01:51 +02:00
jschoubben 2b5119ecd2 Issue 114 is answered: the controller is a process, by ADR 0142
It asked a question rather than reporting a defect, and the question was taken
five days later — the mesh's own components are binaries on the machine, and
third-party software stays a container because an image is the right way to
carry somebody else's build. The delivery of them is issue 142.
2026-09-29 22:38:11 +02:00
jschoubben 96bdffa9bc Two records were numbered 127; the second becomes 149, and is resolved
A number identifies a record, and two were given 127 on 2026-09-27. The one
three documents and three source files cite by number keeps it; the other
becomes 149, says so in its own heading, and its two inbound references are
repointed.

It is also resolved: an empty declaration is sent carrying owns_nothing rather
than skipped, and the host refuses an empty body that does not carry it, so
emptiness cannot be read as a truncated declaration.
2026-09-29 22:33:40 +02:00
jschoubben 4eb16f1028 Three proposed records were already built; two are still yours to call
0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one
assignment of a module per node) described arrangements the mesh has — a
catalogue of other people's software plus a table filled by module add, a vault
answering a secret requirement six modules make, and a rule the assignment
table's primary key already enforces. Each is accepted against what was built,
and says so in its own words.

0037's other half is not built: a manifest outside this catalogue has no check,
which is issue 148.

0068 (the lab takes requests) and 0114 (a shared credential rotates over two
credentials) stay proposed. Neither is built, and both are decisions rather
than records of something that happened.
2026-09-29 22:31:55 +02:00
jschoubben d199de40db Merge the 146/147 records, which 006's note links to 2026-09-29 22:31:46 +02:00
jschoubben 9a1dc4665c Grooming: issue 006's knowledge base is the predecessor's, and is gone
The record catches README claiming this repository is indexed into a knowledge
base. Nothing indexed it, and since the cut-over there is nothing to index it
into — the surface that answered is on the transport the mesh removed (issue
147). Noted where the record is, so the next reader does not go looking for a
search that cannot exist.
2026-09-29 22:26:13 +02:00
jschoubben 14be8576f8 Grooming: five issues were fixed and never closed, and one is not
088, 089, 120, 128 and 130 each name a commit that is on main and cites them —
the forge's address following a moved port, a route naming its endpoint, a
provisioner asking the backend what is there, the hosts file written into a
marked block, and undeclaring giving a unit back the state it was found in.
Each says it was closed by reading commits rather than by a run, so nobody
reads a green that was never measured.

129 stays located on purpose: ca-trust is merged and no machine holds it, so
the symptom it opened on is still true everywhere.
2026-09-29 22:25:25 +02:00
jschoubben ec8676c225 Issue 140 is resolved, and was resolved before it was read again
Written 2026-09-28, answered the same week by mesh-controller bdf965d and
c68d3a7, and never closed — so the mesh's own account said reach was declared
nowhere while three readers were reading it: the filter, the proxy's names, and
each authority's host policy. The resolution names them and what checks each.

What is left is retiring the older per-port keys the block replaces, which is
not a gap in what reach can say.
2026-09-29 22:19:24 +02:00