insomnia.club back to site
[ blog · buyer guide ]

How to scope an AI project: the eight-step worksheet I use before quoting anything

Most AI projects that go wrong were scoped wrong. The model was fine. The team picked a task nobody needed automated, or skipped the data check, or never agreed on what correct means. This is the worksheet I work through with a client before I put a number on anything. You can do most of it yourself, and if you do, any builder you talk to will quote faster and more accurately.

the short version

  • Start from a task inventory with volumes and costs, not from a list of AI ideas.
  • The cost of a wrong answer decides how much review, testing, and design the system needs, and therefore most of the budget.
  • Look at the actual data before you scope. Its condition changes the plan more than any other single input.
  • Write the success metric and build the evaluation set during scoping, so done is defined before work begins.

01why scoping is where AI projects are won

Software scoping has always mattered. AI makes it matter more for one reason: the output is probabilistic. A traditional feature either works or it does not. An AI feature works some percentage of the time on some kinds of inputs, and the scope has to say which percentage, on which inputs, and what happens the rest of the time. If that is not decided before the build, it gets argued about after the demo.

Here is the worksheet. Each step has a short output you can write down. By the end you have a scoping document a builder can price against.

02step 1: task inventory

List the tasks, not the ideas. "Use AI for customer service" is an idea. "Read inbound support emails, classify them, and draft a reply for the agent to approve" is a task.

For each candidate task, write down who does it today, the input (an email, a form, a PDF, a call), the output (a reply, a record, a decision), and the steps in between. Talk to the person who actually does the work. Managers describe the process as designed. Operators describe it as it really runs.

Output: a list of specific tasks, each with input, output, owner, and steps.

03step 2: volume and current cost

For each task, estimate how often it happens and how long it takes. Multiply. That is the time cost. Add anything else it costs: delays that lose customers, overtime, outsourced help.

You do not need precision. You need enough to rank tasks against each other. A task that happens a few times a month is rarely worth custom software, however tedious it is.

Output: each task with rough frequency, time per instance, and total cost.

04step 3: cost of a wrong answer

This is the step most teams skip, and it drives more of the budget than anything else. Ask: if the AI gets this wrong, what happens?

Error costExampleDesign implication
LowAn internal summary that is slightly offLight review, act directly
MediumA drafted customer reply with the wrong toneHuman approves before sending
HighA wrong price, a missed clinical flag, a misfiled legal deadlineStrict review, confidence thresholds, audit trail, narrow automation

High error cost does not rule a task out. It means the system assists a person rather than replacing the decision, and the review interface becomes a real part of the build.

Output: error cost rating and review model for each task.

05step 4: data audit

Pull real samples. Twenty to fifty examples of the inputs and outputs for the task, taken from your actual systems, not idealized. Then answer:

If the AI needs to answer from your documents and records, the build will likely include retrieval, often called RAG, and the quality of that layer depends entirely on what you find here.

Output: data sources, condition, access rules, and constraints per task.

06step 5: success metric

Pick one primary number per task that tells you whether the system earned its keep. Hours saved per week, turnaround time, error rate, conversion, revenue per customer. Then write down how you will measure it, and measure the baseline now, before anything changes.

Avoid metrics you cannot observe. "Better decisions" is not a metric. "Share of drafted replies sent without edits" is.

Output: metric, baseline, target, and measurement method.

07step 6: evaluation set

From your data samples, build a set of real examples with the correct output for each. This is the test the system has to pass. It turns "is it accurate?" into a number and lets you catch hallucinations and regressions every time a prompt or model version changes.

Include the hard cases on purpose: incomplete inputs, unusual requests, edge cases your best operator handles by instinct. A test set made only of easy examples will tell you the system is ready when it is not.

Agree on the pass threshold now. It becomes part of the acceptance criteria in a fixed-budget scope, which I explain in what a fixed-budget software project includes.

Output: evaluation set and agreed pass threshold.

08step 7: integration map

Draw every system the task touches and the direction data moves. Reads are easier than writes. Systems with good APIs are easier than systems you can only reach through exports or screens. Every arrow is work: authentication, error handling, retries, and testing.

Note where a person approves or intervenes, and where the output finally lands. An AI result that sits in a separate tool nobody opens has no value. If the task involves an AI agent taking actions, mark exactly which actions it may take on its own and which require approval.

Output: a diagram of systems, data flows, approval points, and permissions.

09step 8: phase plan

Break the build into phases that each deliver working software. A typical shape:

  1. Read-only assist. The system drafts or recommends; a person does everything else. Fast to ship, low risk, and it generates real feedback.
  2. Integrated workflow. The output lands where people work, with review screens and logging.
  3. Selective automation. Cases above a confidence threshold proceed without review; the rest still go to a person.

Each phase should be measurable against the success metric before the next one starts. Put working software in front of users every two weeks within each phase.

Output: phases, what ships in each, and the measurement gate between them.

10ranking the tasks

With the worksheet done, rank tasks by value against difficulty. The best first project is usually high in volume, medium in error cost, built on data that is in decent shape, and touching no more than two or three systems. Save the high-stakes, messy-data tasks for later, once you have a working evaluation habit and a team that trusts the system.

11where Insomnia Club fits

This worksheet is what our diagnose step looks like. We start with the P&L, work through these eight steps with your operators, and turn the result into a fixed budget before any build starts. The build itself might be AI workflow automation, an AI agent, or AI implementation across several systems.

We are not the right partner if you want a broad AI readiness assessment across every department, or a scoping exercise you plan to shop to the lowest bidder. You are welcome to use this worksheet for that, though. It is yours.

common questions

How do you scope an AI project?

List the tasks you want to change and how often they happen, estimate what they cost today and what an error costs, audit the data they depend on, define a measurable success metric, build a set of real examples to test against, map every system involved, and break the build into phases that each deliver working software.

How long should AI project scoping take?

For a single workflow, days to a couple of weeks depending on how quickly the right people and sample data are available. If scoping stretches into months, it has usually turned into a strategy exercise instead of a plan for something specific.

What is an evaluation set in an AI project?

A collection of real examples from your business, each paired with what a correct output looks like. The system is measured against it before launch and after every change, which turns accuracy from an opinion into a number you can track.

What is the most common AI scoping mistake?

Skipping the data audit. Teams assume the information the AI needs is clean and accessible, then discover mid-build that it lives in scanned files, personal inboxes, or someone's head. Looking at real samples during scoping avoids that.

Should the first AI project be small?

Small enough to ship and measure within weeks, valuable enough that people care whether it works. A tiny project nobody uses teaches nothing, and a sprawling one never reaches production.

keep reading

tell us what keeps you up at night.

Scoped by the people who ship it. Priced before we start.

book a call drop your number

info@insomnia.club