Files
hq/02-DECISIONS/0003-agents-are-persistent-employees.md
jschoubben 066f14b5f8 A node is a conversation, and that is not the employee model
Moving this out of 0003 and out of its vocabulary. I had spent three attempts
fitting the node's own session into the agent-as-employee record, each time
bending hired, draining, reassigned and retired to cover something none of them
describe. 0003 is back to its original text.

It belongs in 0004, under what a node is, because that is what it is -- not a
program installed on a node but part of the node. It holds one session
permanently, anything in the mesh can message it, and it remembers across
callers and across weeks. Its system prompt is the engram, which is recorded
here for the first time despite running on every node.

Also recorded: it has its own narrower tool list, so it can go and look rather
than only report about itself; there is no authorisation between nodes, because
every node is the operator's own; and how a node passes a question on is its
own business rather than a protocol field.

Switched off it still answers, and that is the point of having an off state
rather than an absent one. A node with nothing there is a silence somebody has
to diagnose. A node that says it is switched off is not. Same rule the host
follows about a service that does not exist.

0001's summary is corrected too: it had one row for "agents", which is the
conflation being complained about. Two rows now. A node's own session and a
hired worker are built from the same parts and run on entirely different terms.

Left standing and NOT resolved here: 0001 says a node does not authenticate to
a model provider, agents do. A node that holds a session does. That is a real
conflict between what is recorded and what runs, and it needs deciding rather
than a fourth reconciliation from me.
2026-08-29 14:17:25 +02:00

76 lines
3.7 KiB
Markdown

---
topic: the mesh
status: accepted
date: 2026-07-12
deciders: jochen
reconstructed: true
---
# 3. An agent is a persistent employee, not an instance of a pool
> Reconstructed after the fact from the evidence cited below.
## Context
Agents were originally a **pool**: a named kind of worker, scaled to some number of
interchangeable instances. Work went to whichever instance was free.
That model has no place to put the things that turn out to matter. An agent that accumulates
knowledge of a domain cannot keep it, because the next task lands on a different instance. An
agent cannot own a workspace, because there are several of it. It cannot be held to a policy —
warned for a violation, then dismissed — because there is no continuing subject to warn.
Scaling was also solving a problem the mesh does not have. Instances were being multiplied to
get concurrency, when concurrency is a property of how much work one agent may hold at once.
## Considered options
1. **Keep the pool, attach memory to the pool.** Rejected: shared memory across
interchangeable workers is a knowledge base, not an agent's experience, and the mesh
already has one.
2. **Keep the pool, make instances sticky.** Rejected as a pool pretending to be identities —
identity by scheduling accident, lost on any restart.
3. **One agent is one persistent identity, with concurrency as a property of it.** Chosen.
## Decision
An agent is a **singular, named, persistent identity**: a home node, a workspace on that node,
accumulating memory, and a lifecycle — hired, active, draining, retired. Not a pool member.
Concurrency is a property of the agent, not a count of copies: an agent has a cap on how many
sessions it may hold at once.
Lifecycle is explicit and has verbs. An agent is hired onto a node; it may be reassigned while
idle; it is retired by draining first, and forced only deliberately. Retired agents are not
deleted.
Surge capacity is expressed within the model rather than against it: a template agent is a
blueprint, cloned into a real agent with a lifetime when a queue grows, drained and retired
when it expires. A temporary employee is still an employee.
Some agents are **human**. What differs is modality — how the agent acts — not category. A node
itself is an agent of a kind exempt from the hiring lifecycle.
## Consequences
- Memory, workspace and reputation have a subject to belong to. Policy becomes possible: an
agent that violates a rule can be warned, and warned agents can be dismissed.
- The mesh gained a hiring model, and with it the question of who may hire.
- Scaling by adding instances is gone. If one agent is saturated, either its session cap rises
or another agent is hired — both deliberate acts.
- The transition was not free. Lifecycle columns had to reach every query that selects an
agent, and the ones that were missed failed at the moment of hiring rather than at startup.
- This is the decision [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) generalises:
one kind of participant, differing only in modality.
## References
- `docs(adr): agents as persistent employees + MINERVA librarian` (#495), 2026-07-12 — the
original record, in the code repository.
- `feat(B4): one persistent employee, N sessions — rename max_instances → max_sessions` (#547)
and `feat(noxflow): B3 — workspace provisioner for agent employee model` (#549), 2026-07-20.
- `feat(noxflow): warn-then-fire agents who merge to main without review` (#209), 2026-06-01 —
policy that presumes a continuing subject, predating the model that provides one.
- Knowledge base: `agents/employee-lifecycle`, `agents/temp-surge`, `agents/workspace-layout`.
- The migration cost: `troubleshooting`/`noxflow-agent-enriched-select-missing-lifecycle-columns` (#546).