Issue 147: the operator's tools still dial the bus that was removed

Every tool call fails with 'AMQP not connected', on every node including the
local one, because the tool surface still opens an AMQP connection and that
transport was deleted at the cut-over. The mesh reports healthy throughout —
what broke is the thing standing outside asking it questions, so nothing the
mesh checks is about it.
This commit is contained in:
2026-09-29 21:39:55 +02:00
parent eef54917ec
commit 0dd00e88b6
@@ -0,0 +1,58 @@
---
status: open
opened: 2026-09-29
located-in: []
---
# 147 — the operator's tools still dial the bus that was removed
## What was observed
Every tool call an operator makes against the mesh fails, on every machine, with the same answer:
```
AMQP not connected — cannot reach hal/mesh@novox
AMQP not connected — cannot reach hal/mesh@shanks
```
The mesh moved to one bus and the previous transport was deleted
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over
2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant
speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers,
including the machine the operator is sitting at.
**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries
the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a
node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on
the machine and running the control plane's binary inside its container, which is precisely the
path the tool surface exists to remove, and which nothing checks, records or permits.
It also silently changes how work gets done: an assistant told to use the mesh's tools finds them
dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than
the fault.
## Why this is here and not a note in the knowledge base
The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is
that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke
is not a module, a node or a provision but the thing standing outside asking them questions.
The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same
reason, which is why lessons from the last two days were written into this repository by hand.
## What would have prevented it
- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather
than a separate bridge with its own connection settings that nothing resolves.
- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check
the mesh makes today is about the relationship between the mesh and a machine; none asks whether
a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
the same shape one level out).
## Evidence to carry into diagnosis
- The failure text names `hal/mesh@<node>`, so the bridge is resolving a node and then dialling the
old transport.
- It fails identically for the local machine, which rules out reachability and points at the
transport alone.
- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.