Files
hq/02-DECISIONS/0003-agents-are-persistent-employees.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

3.7 KiB

status, date, deciders, reconstructed
status date deciders reconstructed
accepted 2026-07-12 jochen true

3. An agent is a persistent employee, not an instance of a pool

Reconstructed after the fact from the evidence cited below.

Context

Agents were originally a pool: a named kind of worker, scaled to some number of interchangeable instances. Work went to whichever instance was free.

That model has no place to put the things that turn out to matter. An agent that accumulates knowledge of a domain cannot keep it, because the next task lands on a different instance. An agent cannot own a workspace, because there are several of it. It cannot be held to a policy — warned for a violation, then dismissed — because there is no continuing subject to warn.

Scaling was also solving a problem the mesh does not have. Instances were being multiplied to get concurrency, when concurrency is a property of how much work one agent may hold at once.

Considered options

  1. Keep the pool, attach memory to the pool. Rejected: shared memory across interchangeable workers is a knowledge base, not an agent's experience, and the mesh already has one.
  2. Keep the pool, make instances sticky. Rejected as a pool pretending to be identities — identity by scheduling accident, lost on any restart.
  3. One agent is one persistent identity, with concurrency as a property of it. Chosen.

Decision

An agent is a singular, named, persistent identity: a home node, a workspace on that node, accumulating memory, and a lifecycle — hired, active, draining, retired. Not a pool member.

Concurrency is a property of the agent, not a count of copies: an agent has a cap on how many sessions it may hold at once.

Lifecycle is explicit and has verbs. An agent is hired onto a node; it may be reassigned while idle; it is retired by draining first, and forced only deliberately. Retired agents are not deleted.

Surge capacity is expressed within the model rather than against it: a template agent is a blueprint, cloned into a real agent with a lifetime when a queue grows, drained and retired when it expires. A temporary employee is still an employee.

Some agents are human. What differs is modality — how the agent acts — not category. A node itself is an agent of a kind exempt from the hiring lifecycle.

Consequences

  • Memory, workspace and reputation have a subject to belong to. Policy becomes possible: an agent that violates a rule can be warned, and warned agents can be dismissed.
  • The mesh gained a hiring model, and with it the question of who may hire.
  • Scaling by adding instances is gone. If one agent is saturated, either its session cap rises or another agent is hired — both deliberate acts.
  • The transition was not free. Lifecycle columns had to reach every query that selects an agent, and the ones that were missed failed at the moment of hiring rather than at startup.
  • This is the decision ADR 0001 generalises: one kind of participant, differing only in modality.

References

  • docs(adr): agents as persistent employees + MINERVA librarian (#495), 2026-07-12 — the original record, in the code repository.
  • feat(B4): one persistent employee, N sessions — rename max_instances → max_sessions (#547) and feat(noxflow): B3 — workspace provisioner for agent employee model (#549), 2026-07-20.
  • feat(noxflow): warn-then-fire agents who merge to main without review (#209), 2026-06-01 — policy that presumes a continuing subject, predating the model that provides one.
  • Knowledge base: agents/employee-lifecycle, agents/temp-surge, agents/workspace-layout.
  • The migration cost: troubleshooting/noxflow-agent-enriched-select-missing-lifecycle-columns (#546).