Commit Graph
941 Commits
Author SHA1 Message Date
jochen 68009b16fe Hear what providers retire, and let a person approve, reject and delete (hq ADR 0230)
A provider now waits for a person before retiring more than three consumers
or half of what it holds, and deletes only when asked. The controller is that
person's way in: it keeps waiting and rejected sets as conditions, answers them
with retire approve|reject, lists and deletes retired consumers through the
provider's own tools on its machine, records each act in the hand-act log, and
probes for anything retired longer than thirty days (D11).
2026-10-06 14:41:15 +02:00
mesh-admin 298daa6fbe Merge pull request 'Heal what is known, under a brake, and say every repair (hq to-be 45 Phase 3)' (#85) from feat/a-core-that-cannot-fail-silently-phase-3 into main 2026-10-06 12:29:15 +00:00
jochen 751e39186c Heal what is known, under a brake, and say every repair (hq to-be 45 Phase 3)
Research 031 counted the repairs people made by hand: a push to unstick a
plan waiting on a report, a controller restarted to make an object again, a
plan closed, a consumer re-made from now. Each was the ordinary path taken
again by someone who noticed. The healer registry makes each a registered
response to one condition kind, with a budget, a settle and its event:

- H1 sent-not-reported: ask the machine's node-engine to report again
  (mesh.node.<n>.ask.report); if it does not report what it was sent, send
  it again, never moving a build a policy or a plan holds back
- H2 stalled: close a plan whose wait is superseded or finished
- H3 holder-silent / consumer-lost: the send's own assertion of the bus's
  objects (issue 208's note)
- H4 consumer-behind: consumer-reset, only for a consumer the stream table
  marks resettable (the controller's own events consumer)
- H5 is the identity provider's own repair (ADR 0224 §5), registered only

Success is the observation clearing the condition, never the healer; a spent
budget hands the condition to the operator, urgent, with what was tried, and
no healer touches it again. Every act is begun in the store before it is made
(migration 0070), kept in the condition's tried as "healer Hn" and said as
the seat event healer-acted; a heal is never a hand act. More than twelve acts
in an hour stop every healer until an hour after the last, said urgently.
Only the lease holder heals.

S15 is live: a cause repaired by hand twice in a fortnight raises
healer-wanted, naming the healer that was not enough where one exists. D6's
far-behind finding has its own kind, consumer-behind. Nodes are granted the
question; the controller's grant gains healer-acted (genesis lock in
mesh-host). `healers` lists the registry, the acts and the brake; status
counts the week's heals.
2026-10-06 14:26:26 +02:00
mesh-admin 20c147ffdb Merge pull request 'Act under a lease, keep accounts by order, one writer at composition (hq to-be 45 Phase 2)' (#83) from feat/a-core-that-cannot-fail-silently-phase-2 into main 2026-10-06 10:30:29 +00:00
jochen 2eb9a22c24 Act under a lease, keep accounts by order, one writer at composition (hq to-be 45 Phase 2)
Two controllers could both act (issue 204), a reconcile's report could
overtake the apply after it and the digest decided (issue 267), and a grant
could make a second writer of a machine's report.

- The lease (internal/lease, ADR 0229): mesh-controller_lease key `holder`,
  15 s age, renewed every 5 s by compare-and-set; the epoch is the revision
  it was taken at. The gate is the clock (stops 3 s before expiry); a refused
  renewal is a loss and the process exits; a holder that stops gives it back.
  serve takes it before asserting the bus. Epochs kept in the store
  (migration 0068 controller_epoch) as a floor: a bucket raised from nothing
  is compacted past it. Unleased (no epoch, S12 urgent) only when nobody
  holds it and the bus will not let it be written. A shell command acts
  under the holder's epoch, or its own lease when none.
- Declarations carry `epoch` inside the signed envelope, only to a machine
  whose latest account carried a report_sequence (mesh-host #35); would-send
  is composed with the epoch last sent. Allot and the send both pass the gate.
- Reports: contract in internal/link/order.go (epoch, sequence,
  report_sequence, older_than, refused_older). Accounts kept by epoch, then
  sequence, then report sequence; older refused, counted; unordered reports
  keep the digest rule. Plans by compare-and-set on a revision, with epoch.
  Conditions and calls carry the epoch and are not written off the lease.
- S12 and S13 (naming the writer by epoch) watched, D5 run; reset of the
  bucket said. Writers table compiled in and enforced in PermissionsFor; the
  controller no longer publishes mesh.control.>. A contract per consumed
  kind, and the empty-on-error lint over the repository.
- mesh-host pinned to its main with the epoch in the validator (D1 validates
  the envelope as sent).

Needs mesh-host's genesis lock with the lease grant (mesh-host PR) for
TestTheInstallersFirstUserListIsWhatTheControllerWouldCompose.
2026-10-06 12:29:18 +02:00
mesh-admin 070ecafc07 Merge pull request 'Replace a value given by hand like one the mesh made (hq ADR 0228)' (#82) from feat/a-given-secret-lives-until-the-first-good-start into main 2026-10-06 10:16:30 +00:00
jochen e51c6a2cb9 Replace a value given by hand like one the mesh made (hq ADR 0228)
A given own secret the module reads at start is held by nobody but that
module, so the mesh need not read it to replace it: secret rotate now
works on it, and a value given through secret accept is replaced on its
own after the module's first good start under the mesh. Only a value an
outside party issues (own-secrets "issued-by": "outside") or one the
module applies stays as given, refused with the reason.
2026-10-06 12:13:48 +02:00
mesh-admin 722682f1c4 Merge pull request 'Assert the bus's objects on every send, not only at start (hq issue 208)' (#81) from fix/issue-208-bus-objects-on-send into main 2026-10-06 09:47:01 +00:00
jochen b853439792 Assert the bus's objects on every send, not only at start (hq issue 208)
A module or seat holder assigned after the controller started was sent
its declaration and found nothing to bind: messenger on novox
("consumer novox_messenger not found", 2026-10-06) and every first
build-agent holder (2026-10-03). assertBusObjects ran only in the start
raise; a push ensured a module's consumer only when a bus credential was
minted, which a module carried by the runtime never is.

The send's grant now runs the same derivation (assertOnSend) before the
memberships, on every push, cascade, plan send and rotation. A failure
is said in the send's output and raised as bus.objects.unasserted, which
the next send that asserts everything clears; the send itself goes on,
because the objects are the mesh's and holding every machine back for
one would turn one fault into all. Module consumers are each tried and
every failure named. The start raise stays as it was.
2026-10-06 11:45:51 +02:00
mesh-admin 9cf47f4429 Merge pull request 'Grant the self-check its ban-list question, say a refusal at once, judge the engine by its delivered version (hq to-be 45 Phase 1)' (#80) from fix/phase-1-d8-grant-and-d10-version into main 2026-10-06 08:45:16 +00:00
jochen 8ddc019cd2 Grant the self-check its ban-list question, say a refusal at once, judge the engine by its delivered version (hq to-be 45 Phase 1)
Live on 2026-10-06, two of the first self-check's findings were its own:

- D8 asked every machine's node-intrusion-prevention.banned, and the
  controller's grant did not name the subject: the bus refused it 24 times
  and D8 timed out after thirty seconds instead of saying so. The verbs the
  self-check asks are named in broker.VerbsTheSelfCheckAsks and granted
  (mesh.seat.<seat>.tool.<verb>.*); each probe declares the seat verbs it
  calls, askSeatTool refuses an undeclared one, and a test over the
  registry fails a probe whose question the controller is not granted.
  AskSeatTool now returns a refused publish at once ("the bus refused…")
  instead of waiting out its timeout; D8 asks the machines in parallel.
- D10 read every machine as behind right after a push: a node-engine says
  its version as the directory it is delivered into, the archive's digest
  (31045596c83a, catalogue versionOf), and D10 compared that with the
  build's commit (1545b00a). It now compares with the versions the
  registered build is delivered as, and a hand-placed engine's commit.
2026-10-06 10:44:38 +02:00
mesh-admin fab6b0059e Merge pull request 'Say when the mesh is wrong: conditions, watchdogs, the bus's advisories, doctor (hq to-be 45 Phase 1)' (#79) from feat/a-core-that-cannot-fail-silently-phase-1 into main 2026-10-06 08:29:31 +00:00
jochen e1f5d4fdf0 Vendor every dependency, so no build fetches the host's validator (hq to-be 45 D1)
The controller imports mesh-host/validate through a replace onto the forge
that holds it, and every build — the build agent's go build in a fresh
toolchain container, the Dockerfile's go mod download — would have fetched
it through the public proxy and checksum database at build time: a merge
breaking main on the network, the class Phase 1 removes. vendor/ is
committed; go builds from it with nothing fetched, and refuses to build
when it and go.mod disagree, so a pin moved without go mod vendor fails at
once. The Dockerfile copies vendor/ and builds with GOPROXY=off.
2026-10-06 10:29:10 +02:00
jochen bb1607e424 Say when the mesh is wrong: conditions, watchdogs, the bus's advisories, doctor (hq to-be 45 Phase 1)
Every one of the 48 core failures of research 031 was found by a person
looking; the mesh's answers carried the fact for whoever asked and told
nobody.

- The condition store (to-be 45 §2): mesh-controller_conditions, one key
  per open condition, written by compare-and-set so a person's silence
  and the watchdogs never lose each other's word; every transition kept
  ninety days in mesh-controller_condition-history and said as the
  seat's events condition-raised / condition-changed / condition-cleared
  (the condition at the top level, with event, at, change, why, show),
  offered again while the bus is away. Raised and cleared by observation
  only; a clearing reopened within ten minutes is the same condition with
  its count up, its silence kept. Verbs: conditions, conditions show,
  conditions silence (a hand act, at most a week), conditions history.
- ADR 0224's provider standing is the first kind, provider-failing, held
  by the provider's events; the provider_standing table is no longer read
  or written (left in place: dropping it is the operator's word).
- status leads with the open conditions, urgent first, and says all well
  only with none open; conditions it cannot read are said and not well.
- The signals table compiled in, one watchdog loop over it every 30s: S1
  heartbeat (3 intervals, asleep machines excepted, control node urgent
  after 30 min), S2 report after a send, S3 plan tier, S4 event loop deaf,
  S5 merge not acted, S6 ask lost, S7 call hung, S8 provider silent, S9
  advisories, S10 self-check silent, S11 node tools silent, S13 stale
  refusals; S12, S14, S15 deferred with their reasons. A row that cannot
  see raises probe-failed and clears nothing. A test generated from the
  table suppresses each signal inside and past its bound.
- The bus's advisories (maximum deliveries, a mesh consumer deleted) and
  the controller's own slow consumer and refused subjects, said in the
  mesh's words.
- doctor: the probe registry D1-D10 (D5 deferred) and DW, every five
  minutes, each in thirty seconds; a probe that cannot run is never a
  pass. D1 validates with mesh-host's own validator. Every run ends with
  the doctor-heartbeat event mesh-watcher listens for.
- The controller is granted its new buckets, events, the two advisories
  and $SRV.INFO; the node tools their tools-alive heartbeat. The streams
  and consumers the controller asserts and the ones D6/D7 expect are one
  derivation.
2026-10-06 10:21:11 +02:00
mesh-admin cf4834a36c Merge pull request 'Keep calls and hand acts on the bus, answer status at once, record durations (hq to-be 45 Phase 0)' (#78) from feat/a-core-that-cannot-fail-silently-phase-0 into main 2026-10-06 07:13:39 +00:00
jochen 9d8cbe7b81 Grant the controller its work queues' cancelled sets (hq issue 269)
A cancel writes the ask's id into the seat's cancelled set before deleting
the ask, and a write is a publish to the bucket's subject, which the
controller was not granted: against a server holding exactly the
controller's composed list, every cancel timed out.
2026-10-06 03:01:19 +02:00
jochen e74c32ed50 Keep calls and hand acts on the bus, answer status at once, record durations (hq to-be 45 Phase 0)
A controller restart lost every call's outcome, `status` composed the mesh
while its caller waited (18.6s live on 2026-10-06, past the 10s window), a
repair by hand left no trace, and the core's bounds had nothing measured to
be set from.

- calls: kept in the controller's bucket mesh-controller_calls (last 1000 or
  14 days, answers bounded to 64 KiB), read by id across a restart; a
  controller starting marks a stopped one's running calls abandoned; each
  call names its caller from the inbox its answer goes to.
- status: the serving controller composes it at start, after news from a
  machine, a build or an acting verb, and every minute; the verb answers the
  last composition at once with when and how long it took. Composing resolves
  each machine once instead of twice.
- hand-act log in mesh-controller_hand-acts: push (required through the seat),
  plans stop/close, broker consumer-reset and the new hand-act record take
  --why/--cause/--condition; `hand-acts` lists them and repeated causes;
  status counts the week's.
- durations (migration 0066): apply (send to first report), heartbeat gap,
  plan tier and build, recorded as heard; `durations` summarises them.
- the controller's seat row takes this binary's definition of its own verbs,
  so the console no longer judges calls against an older build's schema.
- the controller is granted its two buckets' subjects.
2026-10-06 02:59:36 +02:00
mesh-admin 146c48fd96 Merge pull request 'Bound a consumer's identity by the provision it requires (hq issue 263, ADR 0225)' (#76) from fix/263-identity-bound-per-provision into main 2026-10-06 00:28:47 +00:00
mesh-admin 1f3abd3e0e Merge pull request 'rotate: narrow a pair credential to one consuming module (hq issue 268)' (#75) from feat/rotate-one-consuming-module into main 2026-10-06 00:25:31 +00:00
jochen 6d620f77c3 Bound a consumer's identity by the provision it requires (hq issue 263)
The one global 20-character bound made every consumer pay an object
store's key length, even for provisions that keep no name, and a single
overflow refused the provider's whole declaration. An offer now states
its own bound (identity: {max, in} or false); unsaid, a provider told its
consumers keeps 20 and one told nothing keeps none. module check judges
every identity on the longest machine name before merge, and a provider
leaves an overflowing consumer out of its grants and composes, with the
consumer named by push, plan and status (ADR 0225).
2026-10-06 02:16:20 +02:00
jochen f8286c063d rotate: narrow a pair credential to one consuming module (hq issue 268)
A machine runs many consumers of one provision, each with its own
credential. When one module leaks its credential, `rotate <provision>
--consumer <machine>` was the narrowest act and replaced every module's
on that machine, restarting all of them. --module (and the verb's
module argument beside provision) rotates only that module's.
2026-10-06 02:13:48 +02:00
mesh-admin e096b4595a Merge pull request 'Keep the account of the sent declaration over an older one (hq issue 267)' (#74) from fix/stale-report-overwrites into main 2026-10-05 23:47:10 +00:00
jochen 09c0c6b367 Keep the account of the sent declaration over an older one (hq issue 267)
The last report stored per node decides whether a release plan moves on,
and it was whichever arrived last. A report about a declaration the mesh
has moved past now records what it says about the machine but leaves the
account of the apply alone, so arrival order cannot undo the newer.
2026-10-06 01:45:35 +02:00
mesh-admin d1fc25f682 Merge pull request 'Catch up on merges the bus announced and never handed over (hq issue 266)' (#73) from fix/missed-merges-are-caught-up into main 2026-10-05 23:33:18 +00:00
mesh-admin eda457f415 Merge pull request 'Answer every seat call within ten seconds and keep what came of it (hq issue 265)' (#72) from fix/a-verb-answers-before-its-caller-gives-up into main 2026-10-05 23:33:11 +00:00
jochen 59f4d486b1 Catch up on merges the bus announced and never handed over (hq issue 266)
The controller acted only on what its events consumer handed it, so a merge
the bus skipped left modules behind with nothing said. The stream is now read
back every five minutes on a single-filter consumer, and any merge that would
still move a module after ten minutes is said and acted on.
2026-10-06 01:29:21 +02:00
jochen 801552c0eb Answer every seat call within ten seconds and keep what came of it (hq issue 265)
A push outlasted the console's 30s wait and, when it sent the bus its
changed user list, the broker's reload forgot the reply it may send:
the push happened and its caller was told it did not answer. Calls now
answer in full or as running with an id, a push answers before it
sends, refused answers are recorded on their call, and 'calls' reads
them back.
2026-10-06 01:14:58 +02:00
mesh-admin ede9bce6ef Merge pull request 'Refuse a verb argument the seat would pass over; a push without a machine says it is the whole mesh (hq issue 244)' (#71) from fix/verb-schemas into main 2026-10-05 22:43:07 +00:00
jochen 0f0028785c Refuse a verb argument the seat would pass over, and say a push is of the whole mesh
A push naming one machine reached the verb without it and pushed every
machine behind (hq issue 244). The controller now refuses any argument a
verb does not declare, any it composed its command line without, and a
switch that is not true or false; a push that names no machine says first
that it is the whole mesh. Tests walk every served verb: no argument is
ever ignored, and every flag of a verb's command, read from the source, is
in its schema or accounted for. plan gains files, push behind, builds and
plans limit.
2026-10-06 00:37:14 +02:00
mesh-admin 6fdcfad8d3 Merge pull request 'Report a provider that keeps failing a consumer in status (hq ADR 0224)' (#70) from feat/a-provider-failing-a-consumer-is-reported into main 2026-10-05 22:20:45 +00:00
jochen 8d9d33ae85 Report a provider that keeps failing a consumer in status (hq ADR 0224)
The identity provider failed every consumer for a day and status called the
mesh well (hq issue 179). The controller now follows every provider's
provisioner.failing/recovered, keeps the newest failing word per provider,
machine and consumer (migration 0065), and status, its JSON and node show
name it until it recovers. Every module that receives contributions is
granted the two events, so no manifest can forget them.
2026-10-06 00:13:23 +02:00
mesh-admin cc25baa563 Merge pull request 'Rename node-hosts-file to node-hostname, and refuse one seat claimed under two names (hq ADR 0223 part 3)' (#69) from hostname-module into main 2026-10-05 22:11:41 +00:00
mesh-admin 68a2ebdcc3 Merge pull request 'Retire node-resolver-config and the seat need only it used (hq ADR 0223 part 2, step 2 of 2)' (#68) from retire-resolv-conf into main 2026-10-05 22:08:03 +00:00
jochen 69b99eec68 Number the hostname seat migration 0064: it merges after the resolver-config one 2026-10-06 00:07:57 +02:00
jochen 09bd0eec4f Number the resolver-config migration 0063: it merges first 2026-10-06 00:07:54 +02:00
mesh-admin 222a38e050 Merge pull request 'Test resolv.conf as the uplink's, and refuse a second writer of a fact's path (hq ADR 0223 part 2, step 1 of 2)' (#67) from resolv-conf-to-uplink into main 2026-10-05 21:57:03 +00:00
jochen ee99a24f77 Rename node-hosts-file to node-hostname, and refuse one seat claimed under two names (hq ADR 0223)
The seat now covers /etc/hostname too. The migration keeps the old name as
an alias so hosts, still assigned while machines move, holds the same seat.
Claims were compared by spelling, so the old and new module would both have
held it on one machine; they are now compared by the seat they resolve to.
2026-10-05 23:43:15 +02:00
jochen f68521da28 Retire node-resolver-config and the seat need it alone used (hq ADR 0223)
The uplink's holder writes /etc/resolv.conf, so the seat that wrote it and
ADR 0220's dependency of it on the uplink have nothing left to say. The
migration deletes the store's row; nothing holds it once resolv-conf is
unassigned everywhere.
2026-10-05 23:40:39 +02:00
jochen 296064c799 Test the resolver file as the uplink's, and refuse a second writer of a fact's path (hq ADR 0223)
The catalogue moves /etc/resolv.conf from resolv-conf to the three uplink
modules. A rendered fact was not compared with other modules' paths, so two
modules could each write the resolver file on one machine, the last winning
every apply; a fact's path now counts as its module's.
2026-10-05 23:39:15 +02:00
mesh-admin df9231c734 Merge pull request 'A mesh seat may be replicated: the resolver held on two machines (hq ADR 0223)' (#65) from feat/the-mesh-has-two-resolvers into main 2026-10-05 20:48:23 +00:00
jochen 843b709b59 The registry-trust test reads the runtime's module, which now writes the trust (ADR 0222) 2026-10-05 22:47:43 +02:00
jochen f506fb34ec Let the mesh's resolver seat have several holders on record
musl takes the first reply from any listed nameserver, so a public fallback
beside the mesh's resolver answered NXDOMAIN for mesh names in every Alpine
container (hq ADR 0223). The fix is two mesh resolvers and no public one, which
needs mesh-dns-resolver held on two machines: a seat can now be replicated,
each holder recorded by 'seat <name> --add', checkClaims accepts every holder
on record and still refuses a second holder of any other mesh seat, a holder
answers its own requirement, and a roster fact gives each replicated seat's
holders, this machine first, so resolv-conf can list them. Migration 0062 keys
a holding by seat and assignment.
2026-10-05 22:42:53 +02:00
mesh-admin c34b937dd3 Merge pull request 'The private network writes nothing into the runtime's file; generated resources are collision-checked (hq ADR 0222, issue 190 — 3 of 3)' (#63) from fix/190-the-overlay-writes-no-runtime-file into main 2026-10-05 20:42:32 +00:00
mesh-admin d84c9699b3 Merge pull request 'Tell a module where a mesh seat's holder is reached: ${seat:<seat>:reach} (hq ADR 0222, issue 190 — 1 of 3)' (#62) from fix/190-seat-reach into main 2026-10-05 20:31:32 +00:00
mesh-admin 002d5e578c Merge pull request 'A named push sends no build a policy or a plan holds back (hq issue 259, ADR 0221)' (#64) from fix/259-a-named-push-sends-no-held-build into main 2026-10-05 20:26:50 +00:00
jochen 2421b82ad2 Keep a named push from sending builds a policy or a plan holds back
A named push flushed every other machine whose declaration differed from
what it was last sent (hq ADR 0083). Under an upgrade policy of `record`,
or a plan still waiting on its first machine (ADR 0218), every machine
running the module differs, so `push <one>` sent the held build to all of
them (hq issue 259).

Each send now records which build of each module it carried
(node.sent_builds, migration 0061). The cascade, and the bus holder added
to a named push, skip a machine any of whose modules would move to a
build its policy records or an open plan has not sent it, and say which
module, which build, why, and that `push <node>` sends it. A machine
whose last send was not recorded is held until it is named. The named
machine itself, a whole-mesh push and `push --behind` are unchanged.
2026-10-05 22:22:03 +02:00
jochen 0ebd48a6a8 The private network writes nothing into the runtime's file (hq issue 190)
daemon.json and docker.service belong to the docker module, which holds node-container-runtime
and now states the registry itself through ${seat:mesh-artifact-store:reach} (hq ADR 0222). The
overlay stops generating registry-trust and registry-trust-reload. A generated resource is now
held to the collision check every module is, so a second writer cannot come back through
computed code; resolution never saw what a generator declares.
2026-10-05 22:19:43 +02:00
jochen 67e291c02a Tell a module where a mesh seat's holder is reached (hq ADR 0222)
The container runtime's module must state the mesh's registry to the runtime it owns, so the
controller can stop writing that into the runtime's file (hq issue 190). ${seat:<seat>:reach}
answers host:port without a binding: nothing required, granted or minted, and the address is
one the mesh already composes into every reference it built. Only mesh-artifact-store is
answered; another seat is refused by name. Unanswered in a file written into as JSON, the empty
member is dropped, so the runtime is never told to trust "".
2026-10-05 22:16:23 +02:00
mesh-admin 8a400d165e Merge pull request 'Delete node-dns-resolver; resolver config needs the uplink (hq ADR 0220)' (#61) from feat/resolver-config-needs-the-uplink into main 2026-10-05 20:11:18 +00:00
jochen 53d3cd7ce9 Delete node-dns-resolver and make resolver config need the uplink
Nothing has claimed node-dns-resolver since the mesh moved to one resolver
(hq ADR 0194); seeding never removes a row, so a migration deletes it.

resolv.conf stays the mesh's only while the network manager is told to keep
off it, which the node-uplink holder does (ADR 0117). A seat's Needs makes
that a dependency checked at assignment by the ADR 0207 mechanism (hq ADR
0220). The two-claimants test keeps its intent with a synthetic module now
that resolved-split-dns leaves the catalogue.
2026-10-05 21:57:04 +02:00