How to choose an AI development company: nine criteria that separate builders from demo shops
Every software shop added AI to its homepage in the last two years. A few of them can ship a system your operators trust at three in the morning. This is the checklist I would use if I were on your side of the table, written by someone who sits on the other side of it every week.
the short version
- Decide what you are buying first: an AI feature in an existing product, a new internal system, or an agent that takes actions. Each needs a different team.
- Judge companies on evaluation plans, integration depth, and who actually ships, not on model names or demo polish.
- Ask for live software and the story of what broke. Teams that ship have scars and can describe them.
- Make sure the code, prompts, test sets, and cloud accounts are in your name from day one.
01first, decide what you are buying
"AI development" covers three very different jobs, and most bad vendor choices start with a buyer who has not decided which one they need.
- An AI feature inside a product you already have. A smarter search box, a summary at the top of a record, a draft reply for your support team. The hard part is fitting into an existing codebase and user experience without breaking either.
- A new internal system. Intake, triage, quoting, reporting, or knowledge lookup that did not exist before. The hard part is the data layer and the workflow design: who sees what, who approves what, and where the output goes.
- An agent that takes actions. Software that reads a situation and then does something in your systems: updates a record, sends a message, books a slot. The hard part is permissions, failure handling, and knowing when to stop and ask a person. (If the term is fuzzy, the glossary entry on AI agents is a good primer.)
Write one paragraph describing which of these you need and what it must return in dollars or hours. That paragraph becomes your filter. A company that pitches you an agent when you asked for a search box is telling you something about how they will run the project.
02the nine criteria
These are the questions that predict whether a project ships and keeps working. For each one I have listed what a good answer sounds like and the red flag to listen for.
1. They start with the business case, not the model
A strong team asks what the work is worth before they ask what model you want. They want to know the volume of the task, what it costs today, and what happens when it goes wrong.
Red flag: the first meeting is a tour of model names and benchmark charts.
2. The people who scope it are the people who ship it
In a lot of agencies, a sales engineer writes the proposal and a different team inherits it. Every assumption that did not make it into the document gets rediscovered at your expense. Ask who will be writing code in week three, and ask to talk to them before you sign.
Red flag: "We will assign the team after kickoff."
3. They can show you working software and tell you what broke
Ask for a live system, not a recording. Then ask the more useful question: what went wrong in production and how did you find out? Teams that ship have specific answers about bad inputs, slow responses, and the time a model update changed behavior. Teams that only demo do not.
Red flag: everything in the portfolio is a prototype or a hackathon project.
4. They have an evaluation plan before they have a prompt
The question that matters most in AI work is "how will we know it is right?" A serious team builds a test set from your real examples, defines what a correct answer looks like, and measures the system against it every time something changes. That is how you control hallucination instead of hoping it away.
Red flag: accuracy is described with adjectives ("very good", "almost always") instead of a method.
5. They are good at data and integration, not just models
Most of the effort in a real system is plumbing: connecting to your CRM, EHR, ERP, or document store, respecting who is allowed to see what, and keeping answers grounded in your records through something like retrieval-augmented generation. Ask how they handle permissions, stale data, and systems with poor APIs.
Red flag: "We will just upload your documents."
6. The pricing puts risk in the right place
Hourly billing on an uncertain project means you carry all of the estimating risk. A fixed budget against a written scope puts it on the builder, which is where it belongs, since they are the ones who know how long things take. What matters next is how changes are handled: priced before they start, or discovered on the invoice. I wrote more about this in what a fixed-budget software project actually includes.
Red flag: a large estimate range with no explanation of what moves it.
7. You see working software early and often
Ask how often you will get something you can click. Every two weeks is a good standard. Long silent build phases are where budgets go to die, because nobody can correct course on software they have not seen.
Red flag: the first real demo is scheduled for month three.
8. They have a plan for after launch
AI systems drift. Model versions change, your data changes, and the edge cases you did not see in testing show up in month two. Ask who monitors quality after launch, how model updates are tested before they reach users, and how usage costs are tracked.
Red flag: "handoff" is the last line of the project plan.
9. You own everything
Code repositories, prompts, evaluation sets, cloud accounts, and model API keys should be in your name from the first day. If the relationship ends, you should be able to hand the whole thing to another team without asking permission.
Red flag: the system runs on the vendor's platform and accounts with no export path.
03a scorecard you can use in the call
| Ask | Good answer | Walk away if |
|---|---|---|
| What is this worth to us? | They ask about volume, cost, and error impact before proposing anything | They quote before understanding the task |
| Who writes the code? | Named people on the call today | A team "to be assigned" |
| How will we know it works? | A test set built from your examples, measured on every change | Adjectives instead of a method |
| What happens when it is wrong? | Confidence thresholds, human review, logging | "It will not be wrong" |
| What do we see in the first month? | Working software in your hands | Documents and wireframes only |
| Who owns the code and accounts? | You, from day one | Their platform, their keys |
04what actually drives the cost
I will not quote ranges here, because a number without a scope is meaningless. These are the factors that move the number on every project I scope:
- Number of systems touched. Each integration adds authentication, error handling, and testing. A system that reads from one database is a different job from one that reads from four and writes to two.
- State of the data. Clean, structured records are cheap to work with. Scanned PDFs, inconsistent spreadsheets, and tribal knowledge are not.
- Accuracy bar. A draft that a person always reviews can tolerate more error than an action taken without review. Higher stakes means more evaluation work.
- Human-in-the-loop design. Review queues, approval steps, and escalation paths are real interface work.
- Compliance. Regulated data such as health records adds hosting, logging, and access control requirements.
- Surfaces. An internal web tool is one build. Native mobile apps for customers are another. Both is both.
- Running costs. Model usage is billed per call. Ask the vendor to estimate it at your real volume, not at demo volume.
05a selection process that takes two weeks, not two quarters
- Write the one-page brief: the job, the volume, what it costs today, what a mistake costs.
- Send the same brief to three companies. Different briefs produce proposals you cannot compare.
- Take one call with each and use the scorecard above. Note who asked the better questions.
- Ask the top one or two for a written scope with a fixed number and a two-week milestone plan.
- Call a reference and ask one thing: what happened when something went wrong?
If the work is still a prototype someone built internally and you need it hardened, the criteria are the same but the scope looks different. See prototype to production for how that engagement is shaped.
06where Insomnia Club fits, honestly
We are a small senior team of operators and engineers. I scope and ship every engagement myself. We work on a fixed budget agreed before work starts, put working software in your hands every two weeks, and stay after launch. Our custom AI development and AI agent work follows exactly the checklist above, because I wrote it from what we do.
Two public examples: Supreme Dental runs an AI smile preview and an LLM assistant inside two native apps that are live in the stores, wired into scheduling and patient records. On Pinned Golf, AI-augmented development reduced the engineering need from five engineers to one.
We are not the right fit if you want engineers billed by the hour under your own management, or if you need fundamental model research. For those, hire contractors or a research lab. If you want a second opinion on a proposal you already have, I am happy to read it. More on how we work is on why Insomnia.
common questions
What does an AI development company actually do?
It designs, builds, and operates software where a model does part of the work: answering questions over your documents, drafting or classifying records, or taking actions inside your systems as an agent. The good ones spend more time on data, integration, and testing than on the model itself.
How much does it cost to hire an AI development company?
It depends on the number of systems the work touches, how clean your data is, how accurate the output has to be, and how many people will use it. Ask for a fixed number tied to a written scope before work starts rather than an open hourly estimate, so the risk of underestimating sits with the builder.
Should I hire an AI development company or freelancers?
Freelancers can work for a narrow, well specified task. A company makes more sense when the work spans data, backend, interface, and ongoing operation, because someone has to own how those pieces fit together and who answers the phone after launch.
How long does an AI development project take?
Scope sets the timeline. A useful sign is how soon you see working software: a team that ships every two weeks will show you something real within the first month, which tells you more than any timeline estimate.
What is the biggest mistake buyers make?
Choosing on demo quality. A demo shows that a model can do the task once. Production requires the system to do it thousands of times, on messy inputs, with a known error rate and a plan for when it is wrong.
