How to evaluate an AI agent vendor: autonomy, permissions, and the exit path
An AI agent is software that decides what to do next and then does it inside your systems. That second half is why buying one is different from buying a chatbot. A chatbot that is wrong says something wrong. An agent that is wrong does something wrong. Here is how I would evaluate a vendor before letting their agent touch anything that matters.
the short version
- Pin down the autonomy level first: does the agent suggest, act with approval, or act on its own? Most problems come from vendors and buyers assuming different answers.
- Every tool the agent can call should be listed, scoped to least privilege, and revocable without a vendor ticket.
- Trial it in a sandbox on your own messy data, scored against a rubric you wrote, not the vendor's demo script.
- Ask about the exit before the entrance: what you can take with you, and what stops working if you leave.
01start by naming the autonomy level
Before features, before pricing, before integrations, get the vendor to say in plain words how much the agent does without a person. I use four levels:
| Level | What the agent does | Good for |
|---|---|---|
| 1. Suggest | Drafts an action; a person performs it | New workflows, high stakes, building trust |
| 2. Act with approval | Prepares the action; a person clicks approve | Customer messages, record changes, anything external |
| 3. Act and report | Acts on its own and logs it for review | Routine, reversible, high-volume work |
| 4. Act silently | Acts on its own with no routine review | Very little, and only after a long track record |
The question to ask: "Can I set the level per action, and change it without your help?" A good answer is yes, with examples. You want to be able to let the agent tag tickets on its own while it still asks before it emails a customer.
If the term itself is fuzzy, the glossary entry on AI agents covers the basics.
02tool permissions: least privilege, written down
An agent is only as dangerous as the tools it can call. Ask for a list of every action the agent can take in your environment, and for each one:
- Which system and which account it uses. Shared admin credentials are a red flag.
- What scope it has: read only, write to specific objects, or broad access.
- Whether it can delete, send externally, or move money. Those deserve their own approval rules.
- How you revoke it. You should be able to cut access from your side in minutes, without opening a support ticket.
Also ask how the agent handles instructions hidden in content it reads. An agent that processes inbound email will eventually read an email that tries to tell it what to do. The vendor should be able to explain how untrusted content is kept separate from instructions, and which actions are off limits no matter what the content says.
03audit logs: can you reconstruct any decision?
When something goes wrong, and it will, you need to answer three questions fast: what did the agent see, what did it decide, and what did it do? Check that the logs capture:
- The input that triggered the run, and any data it retrieved.
- Each step of reasoning or planning the system exposes, and each tool call with its arguments.
- The result of each call, including failures and retries.
- Who approved what, and when.
- The model and configuration version in use at the time.
Then ask where the logs live, how long they are kept, and whether you can export them to your own storage. Logs you can only view inside the vendor's dashboard are logs you lose when you leave.
04human approval: designed, not bolted on
Approval steps are interface work, and vendors vary a lot in how seriously they take it. Look for:
- Approvals that show the person enough context to decide quickly: what the agent proposes, why, and what it looked at.
- Approval in the tools your team already uses, not only in a separate console nobody opens.
- Edit before approve, not just yes or no.
- Escalation when nobody responds, so actions do not silently pile up.
- Reporting on approval rates. If people approve nearly everything unchanged, that action may be ready for more autonomy. If they edit most of them, it is not.
Managing agents well is a skill your team will need to learn. I wrote about that in turning employees into agent managers.
05run a sandbox trial on your own data
The demo will use clean data and a happy path. Your business does not have either. Insist on a trial structured like this:
- Write the rubric first. What counts as a correct outcome, what counts as a harmful one, and the threshold for moving forward. Get the people who do the work today to sign off on it.
- Pick real cases. Pull a sample of recent work including the strange ones: the angry customer, the incomplete form, the request that needed a judgment call.
- Run in a sandbox or shadow mode. The agent works on copies or proposes actions without executing them, so mistakes cost nothing.
- Score blind where you can. Have reviewers grade agent output and human output without knowing which is which.
- Count the failures by type. One harmful action matters more than ten slightly awkward drafts.
A vendor that will not support a trial on your data is asking you to buy on faith.
06integration: open standards beat custom connectors
Ask how the agent connects to systems the vendor does not already support. The answers fall into three groups: they build a connector for you (slow, and they own it), you write against their proprietary API (you own it, but it only works with them), or they support an open standard such as MCP, the Model Context Protocol. With MCP, a connection to your system can be written once and used by any compatible agent, which matters both for speed and for leaving later. If you need connectors for internal systems, that is the kind of thing our MCP server development work covers.
07the exit path
Ask these before you sign, while you still have leverage:
- What do we own? Prompts, workflow definitions, evaluation sets, logs, fine-tuned models if any.
- In what format can we export them?
- What stops working the day we leave, and what keeps running?
- Are the integrations ours or yours?
- How much notice do you give before changing the underlying model or pricing model?
None of this means you should not buy. It means you should know the cost of switching before it becomes the reason you cannot.
08a one-page scorecard
| Area | Pass | Fail |
|---|---|---|
| Autonomy | Set per action, by you | One global switch, or vendor-controlled |
| Permissions | Listed, scoped, revocable by you | Admin credentials, revocation by ticket |
| Audit | Full trace, exportable | Summary only, dashboard only |
| Approvals | In your tools, with edit and escalation | Separate console, approve or reject only |
| Trial | Your data, your rubric, shadow mode | Their demo, their script |
| Integration | Open standards supported | Proprietary connectors only |
| Exit | Documented export, you own the artifacts | Vague answer |
09where Insomnia Club fits
We are not an agent platform, so we have no product to push in this decision. Sometimes a vendor is the right answer and I will say so. When the workflow is specific to how you run, or touches systems no vendor supports well, we build custom AI agents that pass every row of the scorecard above, with the code and accounts in your name, on a fixed budget agreed before work starts. When agents need to move data between your existing tools, that is often AI workflow automation rather than a full agent, and the smaller build is the better one.
common questions
What should I look for in an AI agent vendor?
A clear autonomy model, least-privilege tool permissions, complete audit logs, human approval for high-stakes actions, a trial on your own data, support for open integration standards such as MCP, and a documented exit path. Demo quality matters far less than those.
Is it safe to let an AI agent take actions in our systems?
It can be, if the actions are scoped narrowly, logged completely, reversible where possible, and gated by human approval wherever a mistake is expensive. Start with the agent suggesting actions, measure how often people accept them, and widen autonomy only when the numbers support it.
What is MCP and why does it matter when buying an agent?
MCP, the Model Context Protocol, is an open standard for connecting AI applications to tools and data sources. A vendor that supports it lets you plug in your own systems through a common interface instead of waiting for them to build a proprietary connector, and makes it easier to switch later.
Should we buy an AI agent platform or build our own?
Buy when the job is common across many companies and a vendor already does it well on data like yours. Build when the workflow is specific to how you operate, touches systems no vendor supports, or is a source of competitive advantage you do not want to rent.
How long should an AI agent trial last?
Long enough to see real variety in your inputs, which for most workflows means a few weeks of shadow or assisted operation rather than a single afternoon demo. Decide the success threshold before the trial starts so the result is not argued after the fact.
