mesh/merge-gate pass: builds mesh-host → ace, g14, novox, shanks; no bus step; every machine composes with the change as it did without (4 of 4 compose)
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery delivered
mesh/delivery-group group feat/a-module-says-how-it-is-healthy delivered: every member is delivered
A container that crash-looped after its compose applied passed every check the gate had: nothing looked at what a module runs. The node-engine now judges every long-running resource on every look — one read of the runtime, one per service manager — keeps the restarts it counts across recreates and its own restarts, and says the state in every report and as an event on change, again every minute while not healthy. It reads only; nothing is restarted for being unhealthy.
93 lines
4.2 KiB
Go
93 lines
4.2 KiB
Go
package link
|
|
|
|
import (
|
|
"context"
|
|
"fmt"
|
|
"time"
|
|
|
|
"github.com/nats-io/nats.go"
|
|
)
|
|
|
|
// Bus is what a host needs of the mesh's bus, in the mesh's own words.
|
|
//
|
|
// A host says exactly two things unprompted: what it applied, and that it is here. They are not
|
|
// the same kind of statement and the difference is the whole of this interface — one must arrive
|
|
// and one must not be insisted on.
|
|
//
|
|
// **The host still imports nothing of the mesh's own** (novox/hq ADR 0005): this is its own
|
|
// interface over its own client libraries, not a contract shared with the controller. The two
|
|
// agree because a conformance fixture holds them to one envelope, which is the only kind of
|
|
// agreement that survives being in different repositories.
|
|
type Bus interface {
|
|
// Report says what this node applied. **It must arrive.** A report that fails leaves the
|
|
// mesh believing the node never answered while the node believes it did, and the two go on
|
|
// disagreeing with nothing anywhere saying so — the shape of fault this project keeps
|
|
// finding. Returns false when it could not be delivered, so the caller can say so.
|
|
Report(ctx context.Context, node string, body []byte) error
|
|
|
|
// Alive says this node is here, and nothing else. **Losing one is nothing**: the next is a
|
|
// minute away and the mesh reads a gap rather than counting arrivals. Insisting on delivery
|
|
// would turn a harmless miss into a logged failure every minute.
|
|
Alive(ctx context.Context, node string, body []byte) error
|
|
}
|
|
|
|
// --- The bus the mesh runs on today -----------------------------------------------------------
|
|
|
|
// OverNATS is the bus as a connection. A report goes through JetStream because it must survive
|
|
// the controller's store restarting; a heartbeat does not, because it must not.
|
|
type OverNATS struct {
|
|
Conn *nats.Conn
|
|
JS nats.JetStreamContext
|
|
}
|
|
|
|
// ReportSubject and AliveSubject are this node's own, and no other node's: a host's account may
|
|
// publish `mesh.control.<its own node>.>` and nothing wider, so the subject is the authority on
|
|
// which node a report is about.
|
|
func ReportSubject(node string) string { return "mesh.control." + node + ".report" }
|
|
func AliveSubject(node string) string { return "mesh.control." + node + ".alive" }
|
|
|
|
// HealthSubject is where this node says its long-running resources' health between reports (novox/hq
|
|
// ADR 0240): its own, inside the grant every host already has, and on core NATS like the heartbeat — a
|
|
// statement lost is said again within a minute while anything is not healthy.
|
|
func HealthSubject(node string) string { return "mesh.control." + node + ".health" }
|
|
|
|
// HealthBus is a link that can say a health statement. Separate from Bus so a link that cannot is still
|
|
// a link: the statement then waits for the next report, which carries it too.
|
|
type HealthBus interface {
|
|
Health(ctx context.Context, node string, body []byte) error
|
|
}
|
|
|
|
// Health says a statement, core and flushed, as a heartbeat is.
|
|
func (b OverNATS) Health(ctx context.Context, node string, body []byte) error {
|
|
if err := b.Conn.Publish(HealthSubject(node), body); err != nil {
|
|
return err
|
|
}
|
|
flush, cancel := context.WithTimeout(ctx, 2*time.Second)
|
|
defer cancel()
|
|
return b.Conn.FlushWithContext(flush)
|
|
}
|
|
|
|
func (b OverNATS) Report(ctx context.Context, node string, body []byte) error {
|
|
// Into the CONTROL stream and awaited: this is the message the store-window guarantee is
|
|
// about (novox/hq ADR 0083). The controller naks with a delay while its store is away and
|
|
// the message is redelivered; a publish the bus never accepted must fail here rather than
|
|
// be assumed.
|
|
if _, err := b.JS.Publish(ReportSubject(node), body, nats.Context(ctx)); err != nil {
|
|
return fmt.Errorf("reporting: %w", err)
|
|
}
|
|
return nil
|
|
}
|
|
|
|
func (b OverNATS) Alive(ctx context.Context, node string, body []byte) error {
|
|
// Core, deliberately: a heartbeat in a stream is the mesh's least valuable message competing
|
|
// for retention with its most valuable, and a lost one is the next one.
|
|
if err := b.Conn.Publish(AliveSubject(node), body); err != nil {
|
|
return err
|
|
}
|
|
// Flushed rather than fired and forgotten, so "could not tell the mesh" means the write
|
|
// failed rather than that nobody has looked yet.
|
|
flush, cancel := context.WithTimeout(ctx, 2*time.Second)
|
|
defer cancel()
|
|
return b.Conn.FlushWithContext(flush)
|
|
}
|