AI proof of concept to production: the checklist I use before anything goes live
A proof of concept answers one question: can a model do this task at all? Production asks a harder one: can it do the task thousands of times, on inputs nobody planned for, with a known error rate, a known cost, and somebody accountable when it slips? This is the checklist I run before I let anything cross that line.
the short version
- A proof of concept proves possibility. Production proves repeatability, cost, and accountability. Most of the work lives in the gap.
- Build the evaluation set before you harden anything. Without it, every later change is a guess.
- Decide what the system does when it is unsure, when a dependency is down, and when it is simply wrong, before users find out for you.
- Roll out to a slice of real traffic with a kill switch, and keep the code, prompts, and accounts in your own name.
01what changes between a demo and a system
The proof of concept usually ran on twenty hand-picked examples, in a notebook or a quick web page, against a model API key on somebody's personal account. It worked, people got excited, and now someone wants it in front of customers or staff next month.
Here is what is different once it is real:
- Inputs. Real users send blurry scans, half-finished sentences, the wrong file, and requests nobody anticipated.
- Volume. A cost that rounds to zero at twenty calls is a budget line at real volume.
- Stakes. A wrong answer in a demo is a laugh. A wrong answer in a patient message or a customer quote is a problem with a name attached.
- Change. Model versions update, your data changes, and the person who wrote the prototype moves on to something else.
Each section below is a group of checks. Anything you cannot tick today is work that needs a scope and a budget. If you have not scoped it yet, start with how to scope an AI project.
02evaluation: know whether it is right
This comes first because every other decision depends on it. If you cannot measure quality, you cannot tell whether a fix helped or hurt.
- A written definition of a correct output, agreed by the people who do the work today.
- An evaluation set built from real examples, including the ugly ones, not just the ones that made the demo look good.
- Known failure categories: wrong facts, missing information, wrong format, wrong tone, refusing when it should answer, answering when it should refuse.
- A repeatable run of the system against the set, with results you can compare release to release.
- A target error rate per category, set by what a mistake costs, not by what feels acceptable.
- A plan for hallucination: grounding in your data, citations back to the source record, and a test that catches invented details.
A useful rule: nobody changes a prompt, a model version, or a retrieval setting without running the evaluation set first. That one habit prevents most production regressions I have seen.
03data: know where answers come from
- Every data source the system reads is listed, with an owner and a refresh schedule.
- Stale data is handled: if a document or record changes, the system's view of it changes too.
- Retrieval is tested on its own. If the system uses retrieval-augmented generation, check that the right records come back before you judge the answer built on them.
- Duplicates, outdated versions, and drafts are excluded or clearly marked.
- You know which data must never be sent to a model provider, and that rule is enforced in code, not in a policy document.
04security and permissions
- The system only sees what the person using it is allowed to see. Retrieval respects the same permissions as your source systems.
- Credentials live in a secrets manager under company accounts, not in code or on a laptop.
- Any action the system can take (write a record, send a message, move money) is listed, and each one has an explicit permission.
- Prompt injection is considered: content from emails, documents, or web pages cannot quietly instruct the system to do something it should not.
- Your vendor agreements cover the data you are sending, including any regulated data such as health records.
05observability: see what it is doing
- Every request is logged with input, retrieved context, output, model version, latency, and cost, with sensitive fields handled according to your retention rules.
- Users have a one-click way to flag a bad output, and flags land somewhere a person reviews.
- Someone looks at a sample of real outputs every week, not just the dashboard.
- Alerts exist for error spikes, latency spikes, and cost spikes.
- Flagged outputs feed back into the evaluation set, so the same mistake is tested forever after.
06cost: know what it costs at real volume
I will not give you numbers here because they depend entirely on your volume and design. What I will tell you is where cost hides:
- Context size. Sending whole documents on every call is the most common reason a bill surprises people. Retrieve what is needed, not everything.
- Retries and chains. A workflow that calls a model five times per request costs five times as much as one that calls it once.
- Model choice per step. Not every step needs the largest model. Classification and routing often run well on smaller, cheaper ones.
- Usage growth. Adoption tends to grow once people trust the tool. Model the cost at the volume you hope for, not the volume you have.
Checklist item: a cost per request, measured, and a monthly estimate at expected volume that finance has seen.
07failure handling: decide before users decide for you
| Situation | What production should do |
|---|---|
| Model is unsure | Say so, route to a person, or ask a clarifying question |
| Model provider is down or slow | Time out cleanly, fall back to a manual path, tell the user |
| Output is the wrong format | Validate before using it, retry once, then escalate |
| Source data is missing | Answer that it cannot find it rather than guessing |
| Action fails halfway | Leave records in a known state and log what happened |
| Output is confidently wrong | Human review on high-stakes paths, user flagging everywhere |
The manual path matters more than people think. If the AI step disappears tomorrow, the business should still run, slower. If it cannot, you have a single point of failure you did not plan for.
08rollout: earn trust in slices
- Shadow mode. The system runs on real inputs but nobody sees the output except the team comparing it to what people actually did.
- Assisted mode. Outputs appear as drafts or suggestions that a person accepts, edits, or rejects. Track those three numbers.
- Partial automation. High-confidence, low-stakes cases flow through. Everything else still goes to a person.
- Wider automation. Only when the measured error rate on real traffic supports it.
At every stage: a kill switch that turns the AI step off without a deploy, and a named person who is allowed to use it.
09ownership: who answers the phone
- Code, prompts, evaluation sets, and infrastructure live in repositories and cloud accounts in your company's name.
- A named owner for quality after launch, with time actually allocated.
- A process for testing new model versions before switching to them.
- Documentation good enough that a new engineer could run the evaluation set on day one.
- Users know who to tell when it is wrong, and they hear back.
10where Insomnia Club fits
A lot of our work starts exactly here: a prototype someone built internally, or one a previous vendor delivered, that works in a demo and nobody trusts with real traffic. Our prototype to production engagement runs this checklist against what you have, prices the gaps as a fixed budget before work starts, and ships the hardening in two-week increments you can test. When the system needs to act inside your tools, that becomes AI agent work with permissions and approvals designed in from the start.
If you want to do it yourself, take this list and do it yourself. It is the same list we use. And if the honest answer is that the prototype should be thrown away and the task done another way, I would rather tell you that in the first call than in month two. More on how we decide is in custom AI development.
common questions
Why do so many AI proofs of concept never reach production?
Because the proof of concept was built to answer whether a model can do the task, and nobody budgeted for the rest: evaluation, integration with real systems, permissions, monitoring, failure handling, and an owner after launch. That work is usually larger than the original prototype.
How long does it take to move an AI prototype to production?
It depends on how many systems the work touches, how clean the data is, and how accurate the output has to be. Scope it against this checklist first. The items you cannot check today are the work, and that list sets the timeline.
Can we reuse the code from our proof of concept?
Often the prompts, the examples, and the lessons are worth more than the code. Prototype code tends to skip error handling, permissions, and logging. Keep what is sound, rewrite what was written to demo, and decide that file by file rather than all or nothing.
What is an evaluation set for an AI system?
A collection of real inputs from your business paired with what a correct output looks like. You run the system against it every time a prompt, model, or data source changes, so you can see whether quality went up or down instead of guessing.
Do we need a human in the loop in production?
For anything where a mistake is expensive or hard to reverse, yes, at least at first. A common pattern is to route low-confidence or high-stakes outputs to a person and let the rest flow through, then widen automation as the measured error rate earns it.
