Files
mesh-controller/internal/inventory/migrations/0011-what-a-node-did.sql
T
jschoubben 9681b288aa Keep what each machine did, so status can say what is wrong
A node reports back after applying a declaration: it worked, some of it
failed, or the whole thing was refused. A refusal or a failure moved
last_seen and the reason went to a log line — so "which machine is not
doing what it was told" had no answer the next morning, which is the
question a mesh exists to answer.

Refused and failed are kept as different things, because they are
different situations with different remedies: refused means the machine
is exactly as it was and what is wrong is in what was sent; failed means
it is in a state nobody declared and what is wrong is on the machine. One
word for both would make the record say less than the node did.

One row per node, replaced. The question is the machine's current state —
"this failed an hour ago and then succeeded" is not a machine anybody
needs to look at, and a table of every report would bury the ones that
matter under the ones that do not.

`status` now answers three questions in the order somebody asks them: is
anything broken, is anything not answering, is anything out of date. The
first has consequences now, the third is a plan for later, and a status
leading with the third would bury the first. A machine that has never
spoken is reported as quiet rather than as broken — new, switched off and
unreachable are not the same as tried and could not.

The mapping from a report to an outcome had no test at all, which the
injection caught: it is the code deciding which of those situations a
machine is in. It has four now, including that a partial report never
becomes the account of what the machine holds — the fault that destroyed
a substrate once.
2026-08-30 18:08:59 +02:00

34 lines
1.7 KiB
SQL

-- What each machine did with what it was last sent.
--
-- A node reports back after applying a declaration: everything worked, some of it failed, or the
-- whole thing was refused. Until this, a refusal or a failure moved `last_seen` and the reason
-- went to a log line -- so **which machine is not doing what it was told** had no answer the next
-- morning, and that is the question a mesh exists to answer.
--
-- One row per node, replaced. The question is about the machine's CURRENT state, not its history:
-- "this failed an hour ago and then succeeded" is not a machine anybody needs to look at, and a
-- table of every report would bury the ones that matter under the ones that do not.
create table node_report (
node uuid primary key references node(id) on delete cascade,
-- applied · failed · refused. Named rather than a boolean, because "some of it failed" and
-- "none of it was applied" are different situations with different remedies, and collapsing
-- them would make the report say less than the node did.
outcome text not null,
-- Why, when the whole declaration was refused. The node's own words: the host says exactly
-- what it could not accept, and anything this end wrote instead would be a second, worse
-- explanation of the same thing.
refused text not null default '',
-- Which resources failed, when some did, as [{id, error}].
failed jsonb not null default '[]',
-- How many were applied, for the ordinary case where nothing is wrong and the only useful
-- fact is that the machine did what it was asked.
applied int not null default 0,
at timestamptz not null default now()
);