Every one of the 48 core failures of research 031 was found by a person
looking; the mesh's answers carried the fact for whoever asked and told
nobody.
- The condition store (to-be 45 §2): mesh-controller_conditions, one key
per open condition, written by compare-and-set so a person's silence
and the watchdogs never lose each other's word; every transition kept
ninety days in mesh-controller_condition-history and said as the
seat's events condition-raised / condition-changed / condition-cleared
(the condition at the top level, with event, at, change, why, show),
offered again while the bus is away. Raised and cleared by observation
only; a clearing reopened within ten minutes is the same condition with
its count up, its silence kept. Verbs: conditions, conditions show,
conditions silence (a hand act, at most a week), conditions history.
- ADR 0224's provider standing is the first kind, provider-failing, held
by the provider's events; the provider_standing table is no longer read
or written (left in place: dropping it is the operator's word).
- status leads with the open conditions, urgent first, and says all well
only with none open; conditions it cannot read are said and not well.
- The signals table compiled in, one watchdog loop over it every 30s: S1
heartbeat (3 intervals, asleep machines excepted, control node urgent
after 30 min), S2 report after a send, S3 plan tier, S4 event loop deaf,
S5 merge not acted, S6 ask lost, S7 call hung, S8 provider silent, S9
advisories, S10 self-check silent, S11 node tools silent, S13 stale
refusals; S12, S14, S15 deferred with their reasons. A row that cannot
see raises probe-failed and clears nothing. A test generated from the
table suppresses each signal inside and past its bound.
- The bus's advisories (maximum deliveries, a mesh consumer deleted) and
the controller's own slow consumer and refused subjects, said in the
mesh's words.
- doctor: the probe registry D1-D10 (D5 deferred) and DW, every five
minutes, each in thirty seconds; a probe that cannot run is never a
pass. D1 validates with mesh-host's own validator. Every run ends with
the doctor-heartbeat event mesh-watcher listens for.
- The controller is granted its new buckets, events, the two advisories
and $SRV.INFO; the node tools their tools-alive heartbeat. The streams
and consumers the controller asserts and the ones D6/D7 expect are one
derivation.
The identity provider failed every consumer for a day and status called the
mesh well (hq issue 179). The controller now follows every provider's
provisioner.failing/recovered, keeps the newest failing word per provider,
machine and consumer (migration 0065), and status, its JSON and node show
name it until it recovers. Every module that receives contributions is
granted the two events, so no manifest can forget them.
The mesh runs on the seat's bus alone (novox/hq ADR 0131, design 28 task 5.5). The old
transport's consume loop, build request, tool ask, management API and account scoping are
deleted, and the bus switch with them; the controller connects to the broker seat and to
nothing else. The store-window tests keep their assertions on a bus-less fake, and the tests
that only made sense for the old transport's in-memory holding go with it.
The controller follows the forge's merges (novox/hq 04-ISSUES/131). For each
module recorded as built from that repository and branch it records the move to
the merge commit and builds it — bases first, because a module built before the
module it stands on is built against the old one and reports success, and a base
that fails stops what stands on it. Nothing is pushed here: what a finished build
does to the machines running the module stays the upgrade's decision.
Two more things the same ordering gives: `build --behind` builds bases first, and
`build --on <module>` rebuilds everything that stands on a module — the rebuild a
changed base needs, which "behind" does not see because their sources did not
move.
The outbound half went behind `Bus` and the transport stopped reaching its
callers; this is the other half, and the larger one. Every handler took
`amqp.Delivery`, so the serving loop could not move to another bus without
moving enrolment, reports, builds, upgrades and catch-up with it in one breath.
`Control` states one message in the mesh's words — took it, dropped it, or held
it for the store — and `Inbound` is where messages come from. The AMQP
implementation is today's loop moved rather than changed: same queues, same
prefetch, same holding, because the mesh is running on it and a bus nothing
speaks yet is no reason to alter the one every node is on.
The window (window.go) is now what decides, instead of the conditions that were
inlined in the loop. Two things that surfaced in the wiring:
**Supersession is asked before the store, not after.** A report about a
declaration the mesh has moved past would otherwise wait out a restarting store
to be written and then overwrite what the node is doing now.
**Half of a report is not about a declaration, and that half is never stale.**
What the machine *is* — the tunnel it took over, the ports its own bundle
holds, what an adopted node found, a node moving its overlay key — reaches the
mesh on a report and nowhere else. A rekey set aside as stale is a node whose
overlay key never moves, and no retry is coming, because the node said it once.
So staleness is asked only of a report that is purely an apply's account.
The one thing holding-in-memory can do that holding-in-the-server cannot is
named rather than hidden: `About` sets aside a held message when a newer one
about the same thing arrives, and the bus being built ignores it because the
digest answers the same question.