Inference cost
Inference cost is what it costs to run an AI model to produce outputs, as opposed to the cost of training it. With hosted models it is billed per token of input and output; with self-hosted models it is the compute, hardware, and operations needed to serve requests. For most companies it is the main recurring AI expense.
01why it matters for a business
Few companies train models from scratch, so the AI line item that grows with usage is inference. It scales with volume, prompt size, answer length, model choice, and how many model calls a single task needs. An agent that makes a dozen calls per task costs very differently from a single classification call, even on the same model.
The levers are well known: route simple work to smaller models, keep prompts lean, cache repeated input where providers support it, use batch processing for work that does not need an instant answer (major providers discount batch requests), and cap the number of steps an agent may take. Leaders should ask for a cost-per-task estimate at real volume before approving a build, and for monitoring that shows actual cost after launch.
02what it looks like in practice
A company classifies inbound emails into a dozen categories, then drafts replies for a few of them. The first build sends every email to the most capable model for both steps. A cost review shows classification works just as well on a small, cheap model, and drafting needs the larger model for only a minority of emails. Nightly reporting moves to a batch API. Cost per email drops substantially with no measurable change in quality, verified by running the same evaluation set before and after.
03common mistakes
- Budgeting from list prices without modeling calls per task and real volume.
- Using the largest model everywhere by default.
- No per-feature cost monitoring, so spend creeps and nobody knows which feature caused it.
- Assuming self-hosting is cheaper. At modest volume, hosted APIs often win once hardware and operations are counted.
04related terms
- TokenA token is the unit of text a language model reads and writes: a whole word, part of a word, a number, or a punctuation mark.
- Total cost of ownership (TCO)Total cost of ownership (TCO) is the full cost of a technology decision over its useful life, not just the purchase price or build quote.
- Open-weight modelAn open-weight model is an AI model whose trained parameters, called weights, are published so anyone can download and run it on their own hardware or cloud, subject to its license.
- ROI of AIThe ROI of AI is the measurable business return from an AI initiative relative to its full cost.
- Reasoning modelA reasoning model is a language model trained to work through a problem step by step before giving its final answer, spending extra computation on intermediate thinking.
05where insomnia club fits
Insomnia Club estimates cost per task during scoping and builds in model routing, caching, and spend monitoring, so the running cost is known before you commit and visible after launch.
see custom AI development →tell us what keeps you up at night.
Scoped by the people who ship it. Priced before we start.
book a call drop your number