Files
hq/03-DESIGN/00-as-is/09-interfaces-and-observability.md
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

3.4 KiB

layer, status, code, updated, decisions
layer status code updated decisions
as-is implemented
hal
2026-08-23
02-DECISIONS/0002-nodes-communicate-over-a-broker.md
02-DECISIONS/0010-delivery.md

Interfaces and observability

How the mesh is reached, and how anyone can tell what it is doing.

Capabilities are the primary interface

The mesh's primary interface is not a web console. It is a set of capabilities, exposed to a session and callable in language.

A capability is contributed by a module and is available on any node, wherever it actually runs: local ones directly, remote ones through a stand-in created at startup that forwards over the broker. The caller does not know the difference, and the credentials never move.

This is the mesh's stated vision made concrete — an agent states an intent and the mesh works out which node holds the thing. It is also why a capability's schema is load-bearing in a way that is easy to underestimate: a parameter name that collides with the transport's own reserved names breaks the call, and a validation-library version mismatch has silently dropped every argument while the call still appeared to succeed.

The board

A web interface presents the mesh — nodes, modules, pipelines, agents, work. It is a view. Its own guidance is that shared logic belongs in the mesh's library rather than inline in the board, precisely so the board does not quietly become a second implementation of the mesh's rules.

Public exposure

Nodes carrying a public name run a reverse proxy as the sole entry point. A module declares the names its interfaces answer on, portably, and the proxy's configuration is generated from those declarations rather than written — generated files are marked as such and anything hand-written beside them is left alone.

Nodes without a public role use a local equivalent with locally-trusted certificates. The generation step degrades quietly on a node with no proxy, which is intended and is one more place where "nothing happened" is the correct outcome and looks identical to a failure.

Certificate issuance currently always targets the public authority's production endpoint, which consumes real quota for every experiment (04-ISSUES/004).

Health

Nodes run checks and report. The mesh's health surface answers whether things are up.

What it does not answer is whether they are correct, and that gap is the recurring theme of this whole system: the deploy path reports transport rather than effect, so absence reads as success. A check that confirms a service is running does not confirm the service is running the code that was just deployed, and a node has been left on old code with a version marker that had already advanced.

Thoughts

Every node's daemon runs a periodic loop that surfaces observations from that node's own context. They are stored in the mesh and can inform a session or trigger action.

It is the one part of the mesh that is not request-driven — the mesh noticing things rather than being asked.

The honest summary

Observability tells you the mesh is up. Establishing that it is right currently means reading the operational record and checking by hand.

That is the gap the lab is designed to close (01-to-be/01-end-to-end-testing.md): a place where a change can be run end to end and a verdict produced, cheaply enough that producing one is routine.