Say when the mesh is wrong: conditions, watchdogs, the bus's advisories, doctor (hq to-be 45 Phase 1)
Every one of the 48 core failures of research 031 was found by a person looking; the mesh's answers carried the fact for whoever asked and told nobody. - The condition store (to-be 45 §2): mesh-controller_conditions, one key per open condition, written by compare-and-set so a person's silence and the watchdogs never lose each other's word; every transition kept ninety days in mesh-controller_condition-history and said as the seat's events condition-raised / condition-changed / condition-cleared (the condition at the top level, with event, at, change, why, show), offered again while the bus is away. Raised and cleared by observation only; a clearing reopened within ten minutes is the same condition with its count up, its silence kept. Verbs: conditions, conditions show, conditions silence (a hand act, at most a week), conditions history. - ADR 0224's provider standing is the first kind, provider-failing, held by the provider's events; the provider_standing table is no longer read or written (left in place: dropping it is the operator's word). - status leads with the open conditions, urgent first, and says all well only with none open; conditions it cannot read are said and not well. - The signals table compiled in, one watchdog loop over it every 30s: S1 heartbeat (3 intervals, asleep machines excepted, control node urgent after 30 min), S2 report after a send, S3 plan tier, S4 event loop deaf, S5 merge not acted, S6 ask lost, S7 call hung, S8 provider silent, S9 advisories, S10 self-check silent, S11 node tools silent, S13 stale refusals; S12, S14, S15 deferred with their reasons. A row that cannot see raises probe-failed and clears nothing. A test generated from the table suppresses each signal inside and past its bound. - The bus's advisories (maximum deliveries, a mesh consumer deleted) and the controller's own slow consumer and refused subjects, said in the mesh's words. - doctor: the probe registry D1-D10 (D5 deferred) and DW, every five minutes, each in thirty seconds; a probe that cannot run is never a pass. D1 validates with mesh-host's own validator. Every run ends with the doctor-heartbeat event mesh-watcher listens for. - The controller is granted its new buckets, events, the two advisories and $SRV.INFO; the node tools their tools-alive heartbeat. The streams and consumers the controller asserts and the ones D6/D7 expect are one derivation.
This commit is contained in:
+124
-77
@@ -2,91 +2,153 @@ package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/novox/mesh-controller/internal/conditions"
|
||||
"github.com/novox/mesh-controller/internal/inventory"
|
||||
"github.com/novox/mesh-controller/internal/link"
|
||||
)
|
||||
|
||||
// A provider that keeps failing a consumer is a problem the controller reports (novox/hq ADR 0224).
|
||||
// A provider that keeps failing a consumer is a problem the controller reports (novox/hq ADR 0224) —
|
||||
// **the condition store's first kind** (to-be 45 §2, ADR 0227).
|
||||
//
|
||||
// On 2026-10-05 the identity provider's provisioner failed every consumer from shortly after midnight
|
||||
// until it was fixed by hand that night — 31,000 refused logins after its database was moved and its
|
||||
// admin kept an older password — and `status` called the mesh well all day (novox/hq issue 179). A
|
||||
// provider now announces a consumer it has failed for minutes; the controller keeps it until the
|
||||
// provider says it recovered; and `status`, its JSON and `node show` name it, breaking "all well".
|
||||
// provider announces a consumer it has failed for minutes; the controller keeps it until the provider
|
||||
// says it recovered; and `status`, its JSON and `node show` name it, breaking "all well". Unchanged in
|
||||
// what it says and when; kept as a condition, `provider.<module>.<node>.<consumer>.failing`, rather
|
||||
// than a row of its own, so it is said outward like every other fault and silenced like one.
|
||||
|
||||
// standings keeps what providers say, in the inventory.
|
||||
type standings struct{ inv *inventory.Inventory }
|
||||
// Kinds of the provider standing.
|
||||
const (
|
||||
kindProviderFailing = "provider-failing"
|
||||
kindProviderSilent = "provider-silent"
|
||||
// sourceProvisioner is what raised a standing: the provider's own event.
|
||||
sourceProvisioner = "provisioner.failing"
|
||||
)
|
||||
|
||||
func (s standings) Stood(ctx context.Context, st link.Standing) (bool, error) {
|
||||
return s.inv.KeepStanding(ctx, st.Failing, inventory.ProviderStanding{
|
||||
Module: st.Module, ProviderNode: st.ProviderNode, Provision: st.Provider,
|
||||
Consumer: st.Consumer, ConsumerNode: st.Node,
|
||||
Class: st.Class, Error: st.Error, Since: st.Since, Attempts: st.Attempts,
|
||||
})
|
||||
// providerSaysAgainWithin is how long a failing word stays current without being said again: twice
|
||||
// the quarter of an hour a provider repeats it at (ADR 0224). Past it, S8.
|
||||
const providerSaysAgainWithin = 30 * time.Minute
|
||||
|
||||
// standings keeps what providers say, as conditions.
|
||||
type standings struct {
|
||||
keeper func() *conditions.Keeper
|
||||
}
|
||||
|
||||
// failingProviders is every consumer a provider still assigned where it ran says it keeps failing.
|
||||
//
|
||||
// **A provider no longer assigned is not asked about.** Its last word stays in the store, and is
|
||||
// not a problem: nothing runs there to fail anybody. Assigned again, its first success for each
|
||||
// consumer clears it.
|
||||
func failingProviders(ctx context.Context, inv *inventory.Inventory) ([]inventory.ProviderStanding, error) {
|
||||
all, err := inv.FailingProviders(ctx)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("what providers say they keep failing cannot be read: %w", err)
|
||||
// standingObservation is a provider's failing word as an observation: the provider, its machine and
|
||||
// the consumer name it, so the same consumer failed again is the same condition.
|
||||
func standingObservation(st link.Standing) conditions.Observation {
|
||||
whom := st.Consumer
|
||||
if st.Node != "" {
|
||||
whom += " on " + st.Node
|
||||
}
|
||||
summary := fmt.Sprintf("%s on %s keeps failing %s: %s, %d attempt(s) since %s", st.Module, st.ProviderNode,
|
||||
whom, orUnclassed(st.Class), st.Attempts, st.Since.UTC().Format("2006-01-02 15:04 MST"))
|
||||
said := orUnclassed(st.Class)
|
||||
if e := firstLine(st.Error); e != "" {
|
||||
said += ": " + e
|
||||
}
|
||||
if st.Provider != "" {
|
||||
said += fmt.Sprintf(" (provision %s, %d attempts)", st.Provider, st.Attempts)
|
||||
}
|
||||
return conditions.Observation{Scope: conditions.ScopeProvider,
|
||||
ID: st.Module + "." + st.ProviderNode + "." + st.Consumer,
|
||||
Token: "failing", Kind: kindProviderFailing, Machine: st.ProviderNode, Also: alsoOn(st.Node, st.ProviderNode),
|
||||
Severity: conditions.Warning, Summary: summary, Said: said, Source: sourceProvisioner}
|
||||
}
|
||||
|
||||
// Stood keeps a provider's newest word: failing raises or observes its condition, recovered clears
|
||||
// it. An error is the store away, and the link holds the message to be asked again — a recovery is
|
||||
// said once, and dropping it would leave a consumer named failing that is fine.
|
||||
func (s standings) Stood(ctx context.Context, st link.Standing) (bool, error) {
|
||||
k := s.keeper()
|
||||
if k == nil {
|
||||
return false, fmt.Errorf("the condition store is not open in this controller: %w", link.ErrTryAgain)
|
||||
}
|
||||
o := standingObservation(st)
|
||||
if !st.Failing {
|
||||
why := "the provider says it recovered"
|
||||
if st.Why != "" {
|
||||
why += ": " + st.Why
|
||||
}
|
||||
cleared, err := k.Clear(ctx, o.Key(), why)
|
||||
return cleared, storeAway(err)
|
||||
}
|
||||
_, err := k.Observe(ctx, o)
|
||||
return false, storeAway(err)
|
||||
}
|
||||
|
||||
// storeAway reads the condition store failing as the bus being away for the moment: the link holds the
|
||||
// message and asks again, as it does for a store restarting (ADR 0083), rather than taking it unkept.
|
||||
func storeAway(err error) error {
|
||||
if err == nil {
|
||||
return nil
|
||||
}
|
||||
return fmt.Errorf("%v: %w", err, link.ErrTryAgain)
|
||||
}
|
||||
|
||||
// providerStandings is every open provider-failing condition, from what is open.
|
||||
func providerStandings(open []conditions.Condition) []conditions.Condition {
|
||||
var out []conditions.Condition
|
||||
for _, c := range open {
|
||||
if c.Kind == kindProviderFailing {
|
||||
out = append(out, c)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// providerOf reads a standing's provider module and machine back from its key.
|
||||
func providerOf(c conditions.Condition) (module, node, consumer string, ok bool) {
|
||||
parts := strings.Split(c.Key, ".")
|
||||
if len(parts) != 5 || parts[0] != conditions.ScopeProvider {
|
||||
return "", "", "", false
|
||||
}
|
||||
return parts[1], parts[2], parts[3], true
|
||||
}
|
||||
|
||||
// unassignedProviders clears the standing of every provider no longer assigned where it ran.
|
||||
//
|
||||
// **A provider no longer assigned is not asked about** (ADR 0224 §4): nothing runs there to fail
|
||||
// anybody, and nothing there will ever say it recovered. The observation that resolves it is the
|
||||
// assignment. Assigned again, its first failure raises it again.
|
||||
func unassignedProviders(ctx context.Context, inv *inventory.Inventory, k *conditions.Keeper,
|
||||
open []conditions.Condition) error {
|
||||
assigned := map[string]map[string]bool{}
|
||||
var out []inventory.ProviderStanding
|
||||
for _, s := range all {
|
||||
on, asked := assigned[s.ProviderNode]
|
||||
for _, c := range providerStandings(open) {
|
||||
module, node, _, ok := providerOf(c)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
on, asked := assigned[node]
|
||||
if !asked {
|
||||
modules, err := inv.Assigned(ctx, s.ProviderNode)
|
||||
modules, err := inv.Assigned(ctx, node)
|
||||
if err != nil {
|
||||
// A provider on a machine the mesh no longer knows has nothing running to fail anybody.
|
||||
modules = nil
|
||||
if errors.Is(err, inventory.ErrNoSuchNode) {
|
||||
modules = nil // a machine the mesh no longer knows runs nothing
|
||||
} else {
|
||||
return fmt.Errorf("what %s is assigned cannot be read: %w", node, err)
|
||||
}
|
||||
}
|
||||
on = map[string]bool{}
|
||||
for _, m := range modules {
|
||||
on[m] = true
|
||||
}
|
||||
assigned[s.ProviderNode] = on
|
||||
assigned[node] = on
|
||||
}
|
||||
if on[s.Module] {
|
||||
out = append(out, s)
|
||||
if !on[module] {
|
||||
if _, err := k.Clear(ctx, c.Key, module+" is no longer assigned to "+node+
|
||||
": nothing runs there to fail anybody"); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// failingLines is how status says them: one consumer per entry, the error under it, and a provider
|
||||
// that stopped repeating itself said so.
|
||||
func failingLines(list []inventory.ProviderStanding, now time.Time) []string {
|
||||
var out []string
|
||||
for _, s := range list {
|
||||
where := s.Module
|
||||
if s.ProviderNode != "" {
|
||||
where += " on " + s.ProviderNode
|
||||
}
|
||||
whom := s.Consumer
|
||||
if s.ConsumerNode != "" {
|
||||
whom += " (" + s.ConsumerNode + ")"
|
||||
}
|
||||
out = append(out, fmt.Sprintf(" %-24s fails %s: %s, for %s (%d attempts since %s)",
|
||||
where, whom, orUnclassed(s.Class), roughly(now.Sub(s.Since)), s.Attempts,
|
||||
s.Since.Local().Format("2006-01-02 15:04")))
|
||||
if e := strings.TrimSpace(s.Error); e != "" {
|
||||
out = append(out, fmt.Sprintf(" %-24s %s", "", firstLine(e)))
|
||||
}
|
||||
if s.Quiet(now) {
|
||||
out = append(out, fmt.Sprintf(" %-24s not said again for %s — the provider has stopped "+
|
||||
"saying anything, so this is its last word", "", roughly(now.Sub(s.SaidAt))))
|
||||
}
|
||||
}
|
||||
return out
|
||||
return nil
|
||||
}
|
||||
|
||||
func orUnclassed(class string) string {
|
||||
@@ -96,25 +158,10 @@ func orUnclassed(class string) string {
|
||||
return class
|
||||
}
|
||||
|
||||
// printFailing is the status section, said when there is anything to say.
|
||||
func printFailing(list []inventory.ProviderStanding, now time.Time) {
|
||||
if len(list) == 0 {
|
||||
return
|
||||
// alsoOn is a consumer's machine, when it is not the provider's.
|
||||
func alsoOn(consumerNode, providerNode string) []string {
|
||||
if consumerNode == "" || consumerNode == providerNode {
|
||||
return nil
|
||||
}
|
||||
fmt.Printf("%d consumer(s) a provider keeps failing (ADR 0224):\n\n", len(list))
|
||||
for _, line := range failingLines(list, now) {
|
||||
fmt.Println(line)
|
||||
}
|
||||
fmt.Printf("\n the provider's journal has every attempt; it says recovered on its next success\n\n")
|
||||
}
|
||||
|
||||
// failingOn is the standings that concern one machine: a provider running there, or a consumer.
|
||||
func failingOn(list []inventory.ProviderStanding, node string) []inventory.ProviderStanding {
|
||||
var out []inventory.ProviderStanding
|
||||
for _, s := range list {
|
||||
if s.ProviderNode == node || s.ConsumerNode == node {
|
||||
out = append(out, s)
|
||||
}
|
||||
}
|
||||
return out
|
||||
return []string{consumerNode}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user