Files
hq/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md
T

6.0 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-29
mesh-catalog modules/mesh-console
mesh-tools src/mesh.ts
mesh-controller internal/broker
ADR 0152; mesh-controller PR 164 (invokes); mesh-tools PR 20 (mesh serve, the tools verb); mesh-catalog PR 181 (mesh-console) 03-DESIGN/01-to-be/34-the-console.md

147 — the operator's tools still dial the bus that was removed

What was observed

Every tool call an operator makes against the mesh fails, on every machine, with the same answer:

AMQP not connected — cannot reach hal/mesh@novox
AMQP not connected — cannot reach hal/mesh@shanks

The mesh moved to one bus and the previous transport was deleted (ADR 0131, cut over 2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers, including the machine the operator is sitting at.

What that leaves. The mesh itself is healthy: nodes are current, modules run, the bus carries the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on the machine and running the control plane's binary inside its container, which is precisely the path the tool surface exists to remove, and which nothing checks, records or permits.

It also silently changes how work gets done: an assistant told to use the mesh's tools finds them dead, falls back to ssh and docker exec, and the operator discovers the fallback rather than the fault.

Why this is here and not a note in the knowledge base

The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke is not a module, a node or a provision but the thing standing outside asking them questions.

The same cut-over removed the assistant's long-term memory (recall, memorize) for the same reason, which is why lessons from the last two days were written into this repository by hand.

What would have prevented it

  • The tool surface as a consumer of the bus like any other, so moving the bus moves it — rather than a separate bridge with its own connection settings that nothing resolves.
  • A check that a tool call reaches a node, run where the mesh's other checks run. Every check the mesh makes today is about the relationship between the mesh and a machine; none asks whether a person can ask it anything (issue 145, the same shape one level out).

Diagnosed at once, because the answer was in the configuration

The surface is not the mesh's. The tool server the operator's assistant speaks to is the predecessor's brain, installed on the workstation and started as a local process, with the predecessor's broker URL — amqp://…@<the control node> — written into the assistant's own configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account on the mesh's bus, and the mesh has never known it exists.

So nothing regressed. The mesh removed a transport that this program still dials, and the program was never part of the mesh to be moved. The mesh has a tool model — tools in a manifest, the calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one — and nothing publishes an operator-facing surface onto it. What an operator uses is the thing that came before, kept alive by a URL in a file.

That is the issue, and it is larger than a broken connection: the way a person drives this mesh is outside the mesh.

Evidence to carry into diagnosis

  • The failure text names hal/mesh@<node>, so the bridge is resolving a node and then dialling the old transport.
  • It fails identically for the local machine, which rules out reachability and points at the transport alone.
  • The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.

Answered, 2026-09-30

The work order's question — does the mesh grow its own operator surface, or is the surface an ordinary module — is answered by ADR 0152: the console is a module. mesh-console is assigned to the machine a person sits at, holds a credential the mesh minted, calls tools under a grant its manifest declares (invokes), and serves the mesh's tools on that machine's loopback to an agent over MCP and to a person through the same endpoint. Designed in 34 — The console.

What that leaves, said plainly so nobody reads this record as closed on the whole of its first paragraph: the console reaches every tool a module serves. The mesh's own questions — what a node runs, what is assigned — are the mesh-controller seat's tools under ADR 0132 and are not on the bus yet; for those a shell is still the way, and design 33 is where that closes.

The predecessor's program on the workstation is not replaced by the mesh; it is left where it is and the assistant is pointed at the console beside it. The hal entry in the assistant's configuration still names things that are not the mesh's.

Verified live, 2026-09-30 evening. The four pull requests merged; the console was registered (checked first with module check), built, assigned to a workstation, issued a bus account, and pushed. On that machine tools/list answered on loopback with 62 tools and named 36 modules as not answering, and a call to the forge's gitea_list_repos returned repositories. The assistant on that machine now lists the console as a connected MCP server beside the predecessor's program, which was left where it is. As-is: 13-the-console.md.