Say when the mesh is wrong: conditions, watchdogs, the bus's advisories, doctor (hq to-be 45 Phase 1)

Every one of the 48 core failures of research 031 was found by a person
looking; the mesh's answers carried the fact for whoever asked and told
nobody.

- The condition store (to-be 45 §2): mesh-controller_conditions, one key
  per open condition, written by compare-and-set so a person's silence
  and the watchdogs never lose each other's word; every transition kept
  ninety days in mesh-controller_condition-history and said as the
  seat's events condition-raised / condition-changed / condition-cleared
  (the condition at the top level, with event, at, change, why, show),
  offered again while the bus is away. Raised and cleared by observation
  only; a clearing reopened within ten minutes is the same condition with
  its count up, its silence kept. Verbs: conditions, conditions show,
  conditions silence (a hand act, at most a week), conditions history.
- ADR 0224's provider standing is the first kind, provider-failing, held
  by the provider's events; the provider_standing table is no longer read
  or written (left in place: dropping it is the operator's word).
- status leads with the open conditions, urgent first, and says all well
  only with none open; conditions it cannot read are said and not well.
- The signals table compiled in, one watchdog loop over it every 30s: S1
  heartbeat (3 intervals, asleep machines excepted, control node urgent
  after 30 min), S2 report after a send, S3 plan tier, S4 event loop deaf,
  S5 merge not acted, S6 ask lost, S7 call hung, S8 provider silent, S9
  advisories, S10 self-check silent, S11 node tools silent, S13 stale
  refusals; S12, S14, S15 deferred with their reasons. A row that cannot
  see raises probe-failed and clears nothing. A test generated from the
  table suppresses each signal inside and past its bound.
- The bus's advisories (maximum deliveries, a mesh consumer deleted) and
  the controller's own slow consumer and refused subjects, said in the
  mesh's words.
- doctor: the probe registry D1-D10 (D5 deferred) and DW, every five
  minutes, each in thirty seconds; a probe that cannot run is never a
  pass. D1 validates with mesh-host's own validator. Every run ends with
  the doctor-heartbeat event mesh-watcher listens for.
- The controller is granted its new buckets, events, the two advisories
  and $SRV.INFO; the node tools their tools-alive heartbeat. The streams
  and consumers the controller asserts and the ones D6/D7 expect are one
  derivation.
This commit is contained in:
jochen
2026-10-06 10:21:11 +02:00
parent cf4834a36c
commit bb1607e424
51 changed files with 6299 additions and 404 deletions
+124 -77
View File
@@ -2,91 +2,153 @@ package main
import (
"context"
"errors"
"fmt"
"strings"
"time"
"github.com/novox/mesh-controller/internal/conditions"
"github.com/novox/mesh-controller/internal/inventory"
"github.com/novox/mesh-controller/internal/link"
)
// A provider that keeps failing a consumer is a problem the controller reports (novox/hq ADR 0224).
// A provider that keeps failing a consumer is a problem the controller reports (novox/hq ADR 0224) —
// **the condition store's first kind** (to-be 45 §2, ADR 0227).
//
// On 2026-10-05 the identity provider's provisioner failed every consumer from shortly after midnight
// until it was fixed by hand that night — 31,000 refused logins after its database was moved and its
// admin kept an older password — and `status` called the mesh well all day (novox/hq issue 179). A
// provider now announces a consumer it has failed for minutes; the controller keeps it until the
// provider says it recovered; and `status`, its JSON and `node show` name it, breaking "all well".
// provider announces a consumer it has failed for minutes; the controller keeps it until the provider
// says it recovered; and `status`, its JSON and `node show` name it, breaking "all well". Unchanged in
// what it says and when; kept as a condition, `provider.<module>.<node>.<consumer>.failing`, rather
// than a row of its own, so it is said outward like every other fault and silenced like one.
// standings keeps what providers say, in the inventory.
type standings struct{ inv *inventory.Inventory }
// Kinds of the provider standing.
const (
kindProviderFailing = "provider-failing"
kindProviderSilent = "provider-silent"
// sourceProvisioner is what raised a standing: the provider's own event.
sourceProvisioner = "provisioner.failing"
)
func (s standings) Stood(ctx context.Context, st link.Standing) (bool, error) {
return s.inv.KeepStanding(ctx, st.Failing, inventory.ProviderStanding{
Module: st.Module, ProviderNode: st.ProviderNode, Provision: st.Provider,
Consumer: st.Consumer, ConsumerNode: st.Node,
Class: st.Class, Error: st.Error, Since: st.Since, Attempts: st.Attempts,
})
// providerSaysAgainWithin is how long a failing word stays current without being said again: twice
// the quarter of an hour a provider repeats it at (ADR 0224). Past it, S8.
const providerSaysAgainWithin = 30 * time.Minute
// standings keeps what providers say, as conditions.
type standings struct {
keeper func() *conditions.Keeper
}
// failingProviders is every consumer a provider still assigned where it ran says it keeps failing.
//
// **A provider no longer assigned is not asked about.** Its last word stays in the store, and is
// not a problem: nothing runs there to fail anybody. Assigned again, its first success for each
// consumer clears it.
func failingProviders(ctx context.Context, inv *inventory.Inventory) ([]inventory.ProviderStanding, error) {
all, err := inv.FailingProviders(ctx)
if err != nil {
return nil, fmt.Errorf("what providers say they keep failing cannot be read: %w", err)
// standingObservation is a provider's failing word as an observation: the provider, its machine and
// the consumer name it, so the same consumer failed again is the same condition.
func standingObservation(st link.Standing) conditions.Observation {
whom := st.Consumer
if st.Node != "" {
whom += " on " + st.Node
}
summary := fmt.Sprintf("%s on %s keeps failing %s: %s, %d attempt(s) since %s", st.Module, st.ProviderNode,
whom, orUnclassed(st.Class), st.Attempts, st.Since.UTC().Format("2006-01-02 15:04 MST"))
said := orUnclassed(st.Class)
if e := firstLine(st.Error); e != "" {
said += ": " + e
}
if st.Provider != "" {
said += fmt.Sprintf(" (provision %s, %d attempts)", st.Provider, st.Attempts)
}
return conditions.Observation{Scope: conditions.ScopeProvider,
ID: st.Module + "." + st.ProviderNode + "." + st.Consumer,
Token: "failing", Kind: kindProviderFailing, Machine: st.ProviderNode, Also: alsoOn(st.Node, st.ProviderNode),
Severity: conditions.Warning, Summary: summary, Said: said, Source: sourceProvisioner}
}
// Stood keeps a provider's newest word: failing raises or observes its condition, recovered clears
// it. An error is the store away, and the link holds the message to be asked again — a recovery is
// said once, and dropping it would leave a consumer named failing that is fine.
func (s standings) Stood(ctx context.Context, st link.Standing) (bool, error) {
k := s.keeper()
if k == nil {
return false, fmt.Errorf("the condition store is not open in this controller: %w", link.ErrTryAgain)
}
o := standingObservation(st)
if !st.Failing {
why := "the provider says it recovered"
if st.Why != "" {
why += ": " + st.Why
}
cleared, err := k.Clear(ctx, o.Key(), why)
return cleared, storeAway(err)
}
_, err := k.Observe(ctx, o)
return false, storeAway(err)
}
// storeAway reads the condition store failing as the bus being away for the moment: the link holds the
// message and asks again, as it does for a store restarting (ADR 0083), rather than taking it unkept.
func storeAway(err error) error {
if err == nil {
return nil
}
return fmt.Errorf("%v: %w", err, link.ErrTryAgain)
}
// providerStandings is every open provider-failing condition, from what is open.
func providerStandings(open []conditions.Condition) []conditions.Condition {
var out []conditions.Condition
for _, c := range open {
if c.Kind == kindProviderFailing {
out = append(out, c)
}
}
return out
}
// providerOf reads a standing's provider module and machine back from its key.
func providerOf(c conditions.Condition) (module, node, consumer string, ok bool) {
parts := strings.Split(c.Key, ".")
if len(parts) != 5 || parts[0] != conditions.ScopeProvider {
return "", "", "", false
}
return parts[1], parts[2], parts[3], true
}
// unassignedProviders clears the standing of every provider no longer assigned where it ran.
//
// **A provider no longer assigned is not asked about** (ADR 0224 §4): nothing runs there to fail
// anybody, and nothing there will ever say it recovered. The observation that resolves it is the
// assignment. Assigned again, its first failure raises it again.
func unassignedProviders(ctx context.Context, inv *inventory.Inventory, k *conditions.Keeper,
open []conditions.Condition) error {
assigned := map[string]map[string]bool{}
var out []inventory.ProviderStanding
for _, s := range all {
on, asked := assigned[s.ProviderNode]
for _, c := range providerStandings(open) {
module, node, _, ok := providerOf(c)
if !ok {
continue
}
on, asked := assigned[node]
if !asked {
modules, err := inv.Assigned(ctx, s.ProviderNode)
modules, err := inv.Assigned(ctx, node)
if err != nil {
// A provider on a machine the mesh no longer knows has nothing running to fail anybody.
modules = nil
if errors.Is(err, inventory.ErrNoSuchNode) {
modules = nil // a machine the mesh no longer knows runs nothing
} else {
return fmt.Errorf("what %s is assigned cannot be read: %w", node, err)
}
}
on = map[string]bool{}
for _, m := range modules {
on[m] = true
}
assigned[s.ProviderNode] = on
assigned[node] = on
}
if on[s.Module] {
out = append(out, s)
if !on[module] {
if _, err := k.Clear(ctx, c.Key, module+" is no longer assigned to "+node+
": nothing runs there to fail anybody"); err != nil {
return err
}
}
}
return out, nil
}
// failingLines is how status says them: one consumer per entry, the error under it, and a provider
// that stopped repeating itself said so.
func failingLines(list []inventory.ProviderStanding, now time.Time) []string {
var out []string
for _, s := range list {
where := s.Module
if s.ProviderNode != "" {
where += " on " + s.ProviderNode
}
whom := s.Consumer
if s.ConsumerNode != "" {
whom += " (" + s.ConsumerNode + ")"
}
out = append(out, fmt.Sprintf(" %-24s fails %s: %s, for %s (%d attempts since %s)",
where, whom, orUnclassed(s.Class), roughly(now.Sub(s.Since)), s.Attempts,
s.Since.Local().Format("2006-01-02 15:04")))
if e := strings.TrimSpace(s.Error); e != "" {
out = append(out, fmt.Sprintf(" %-24s %s", "", firstLine(e)))
}
if s.Quiet(now) {
out = append(out, fmt.Sprintf(" %-24s not said again for %s — the provider has stopped "+
"saying anything, so this is its last word", "", roughly(now.Sub(s.SaidAt))))
}
}
return out
return nil
}
func orUnclassed(class string) string {
@@ -96,25 +158,10 @@ func orUnclassed(class string) string {
return class
}
// printFailing is the status section, said when there is anything to say.
func printFailing(list []inventory.ProviderStanding, now time.Time) {
if len(list) == 0 {
return
// alsoOn is a consumer's machine, when it is not the provider's.
func alsoOn(consumerNode, providerNode string) []string {
if consumerNode == "" || consumerNode == providerNode {
return nil
}
fmt.Printf("%d consumer(s) a provider keeps failing (ADR 0224):\n\n", len(list))
for _, line := range failingLines(list, now) {
fmt.Println(line)
}
fmt.Printf("\n the provider's journal has every attempt; it says recovered on its next success\n\n")
}
// failingOn is the standings that concern one machine: a provider running there, or a consumer.
func failingOn(list []inventory.ProviderStanding, node string) []inventory.ProviderStanding {
var out []inventory.ProviderStanding
for _, s := range list {
if s.ProviderNode == node || s.ConsumerNode == node {
out = append(out, s)
}
}
return out
return []string{consumerNode}
}