Managing AI agents: how to turn your employees into agent managers
The interesting change in companies that use AI well is not that work disappears. It is that the people who used to do a task now supervise something that does it. That is a management job, and almost nobody has been trained for it. This is how we run it inside Insomnia Club and how I would set it up in yours.
the short version
- An agent manager owns outcomes, not keystrokes: they write the task spec, review the output, handle escalations, and improve the agent over time.
- The best agent managers are the people who did the task well by hand. Their judgment is the asset; the agent is leverage on it.
- Every agent needs a written spec, a review checklist, escalation rules, and two or three metrics, the same way a new hire needs a role description.
- Change management decides whether this works. Be explicit about what happens to roles, and give people a real career path as agent managers.
01the shift: from doing to supervising
When an AI agent takes over a task, the work does not vanish. It changes shape. Someone still decides what the task is, checks that it was done right, deals with the cases the agent cannot handle, and makes the agent better next month than it was this month. That someone is an agent manager.
Inside Insomnia Club this is not theory. We run agents for QA and development work every day, and our engineers spend a growing share of their time specifying, reviewing, and correcting agent output rather than typing every line themselves. It is a big part of how a small senior team covers ground that used to need a larger one. On Pinned Golf, AI-augmented development reduced the engineering need from five engineers to one, with the same review discipline a bigger team would apply. The review discipline is the point. It is what made the leverage safe.
The same pattern applies outside engineering: in support, operations, finance, and sales. The details differ; the management job is the same.
02what an agent manager actually does
Think of it as managing a very fast, very literal junior employee who never gets tired and never learns on their own unless you change their instructions. The job has five parts:
- Specify the work. Write down what the agent does, with what inputs, to what standard.
- Review the output. Every item at first, then a sample once quality is proven.
- Handle escalations. Resolve the cases the agent flags as uncertain or out of scope.
- Track the numbers. Volume, acceptance rate, errors, time, cost.
- Improve the agent. Update instructions, add examples, fix tool access, and retire tasks the agent should not be doing.
03the task spec
Most agent failures I see are specification failures. The agent did what it was told; it was told something vague. A good spec answers these questions in writing:
- Goal. What is the finished output, and who uses it?
- Inputs. Which systems and documents can the agent read? Which can it change?
- Standard. Three to five examples of excellent output and two of unacceptable output, with notes on why.
- Boundaries. What the agent must never do: send to a customer without review, change a price, delete a record, merge code to the main branch.
- Definition of done. How the agent knows it finished, and what it hands to the person reviewing.
If you cannot write the spec, the task is not ready for an agent. That is useful information too. Writing it often exposes that two people on your team have been doing the same task two different ways.
04review checklists
A reviewer without a checklist reviews whatever catches their eye. A checklist makes review fast and consistent, and it doubles as the test set you use when you change the agent. For a support reply, it might be: correct customer and order, policy applied correctly, no promise we cannot keep, tone matches our guide, next step clear. For a code change in our own work, it is: tests pass, the change does what the ticket asked and nothing else, no new security exposure, readable by the next person.
Keep the checklist short enough to use on every item. Five to eight checks is usually right. If reviewers start skipping it, it is too long.
05escalation rules
An agent that never escalates is either doing trivial work or hiding mistakes. Define in advance when the agent stops and asks a person:
- Its confidence is low or the inputs conflict.
- The case matches a known sensitive category: a complaint, a legal mention, a large refund, a production database.
- The action is irreversible.
- The request falls outside the spec.
Then decide who receives each escalation and how fast they respond. An escalation that sits for two days is worse than no agent at all, because the customer or colleague waiting on it assumed it was handled.
06metrics per agent
| Metric | What it tells you | Watch for |
|---|---|---|
| Volume handled | How much work moved off people | Volume rising while quality drops |
| Accepted without edits | Whether the spec is right | A plateau that means the spec needs work |
| Escalation rate | Whether boundaries are well set | Near zero (hiding) or very high (wrong task) |
| Errors found in review | Real quality | Errors of the same kind repeating |
| Cycle time | Speed to finished output, review included | Review becoming the bottleneck |
| Cost per completed task | The economics at scale | Model usage creeping up unnoticed |
Review these weekly for a new agent and monthly once it is stable. Changes to the model, the instructions, or the tools should always be followed by a check against your review checklist examples before they reach real work.
07career paths
If agent management is treated as a side duty, your best people will avoid it. Make it a real role with a real ladder:
- Practitioner with agents. Does their job with one or two agents doing parts of it, and reviews the output.
- Agent manager. Owns one or more agents for a function: specs, review, metrics, improvements.
- Agent lead. Owns agent operations across a department, decides which tasks move to agents, and trains new agent managers.
Pay and title should follow scope of outcomes owned, not headcount managed. A person whose agents handle the work of a whole team is doing a more valuable job than they were before, and they should see that reflected.
08change management, honestly
People are not afraid of AI tools. They are afraid of not knowing what happens to them. Three things help:
- Say what is changing and what is not, early and specifically. Vague reassurance is heard as bad news.
- Start with the people who are best at the task. Let them design the spec and the checklist. They become the proof that this is a promotion of their judgment, not a replacement of it.
- Train for the new job. Writing specs and reviewing output are skills. Teach them directly instead of hoping people pick them up.
Leaders go first. If you have not used these tools on your own work, start with how executives should learn AI in 30 days; you will make better decisions about everyone else's role once you have.
09how we help
For engineering teams, AI coding orchestration is the setup we run ourselves: agents doing QA and development tasks under engineers who specify, review, and merge. For other functions, we build the agents through AI agent development and teach your people to manage them through AI training for teams. The spec, checklist, escalation rules, and metrics above are part of every handover, because an agent nobody knows how to manage is a liability, not an asset.
common questions
What does it mean to manage AI agents?
It means owning the outcome of work an AI agent performs: defining the task clearly, reviewing a sample or all of the output, handling the cases the agent escalates, tracking quality and volume, and improving the instructions and tools when the agent falls short. It looks more like managing a fast junior employee than operating software.
Who should become an AI agent manager?
Usually the person who did the task best by hand. They know what good output looks like, which edge cases matter, and when something is off. Technical skill helps but is secondary to judgment about the work.
How many agents can one person manage?
It depends on how much review the work needs. High-stakes output that needs a person to check every item limits one manager to fewer agents. Lower-stakes work with sampled review lets one person oversee much more. Start with full review and widen only when the metrics justify it.
Will AI agents replace employees?
Some tasks will move from people to agents. The companies that handle it well redeploy their best people into managing and improving those agents and into work that needs human judgment. Be direct with your team about what is changing; uncertainty does more damage than the change itself.
What metrics should we track for an AI agent?
Track volume handled, the share of output accepted without edits, the escalation rate, error rate found in review, and the time from task start to finished output. Add cost per completed task so you know the economics at scale.
