Token
A token is the unit of text a language model reads and writes: a whole word, part of a word, a number, or a punctuation mark. Model limits and provider pricing are both measured in tokens. A common rule of thumb for English is that one token averages about three quarters of a word, though it varies by model.
01why it matters for a business
Tokens are the meter on every AI system you run. Providers charge per token, usually quoting a price per million tokens, with input (what you send) and output (what the model writes) priced separately and output typically costing more. A design that stuffs entire documents into every request, or asks for long answers when short ones would do, multiplies cost at scale.
Tokens also define limits. A model's context window is measured in tokens, and so are most provider rate limits. Languages other than English, code, and tables often use more tokens for the same content, which matters if you operate across markets or process structured data.
02what it looks like in practice
A support team summarizes every closed ticket. The first version sends the full thread, all internal notes, and a long instruction block to a large model for each summary. A revised version strips signatures and quoted email history, sends only the fields that matter, keeps the instructions short and reusable, and routes routine tickets to a smaller model. The output is the same, the tokens per summary fall sharply, and the monthly bill follows. None of that required a new vendor, only attention to what goes into each request.
03common mistakes
- Estimating cost from a demo. Multiply tokens per task by real monthly volume before committing.
- Forgetting output tokens. Verbose answers cost more than terse ones, so ask for the format you need.
- Ignoring caching. Several providers discount repeated input, such as a long fixed instruction block.
- Treating word counts and token counts as the same when planning limits.
04related terms
- Inference costInference cost is what it costs to run an AI model to produce outputs, as opposed to the cost of training it.
- Context windowA context window is the maximum amount of text, measured in tokens, that a language model can take into account in a single request, including your instructions, any documents or conversation history, and the answer it writes.
- Large language model (LLM)A large language model (LLM) is an AI model trained on very large amounts of text to predict the next piece of text, which lets it write, summarize, translate, classify, extract information, and reason through problems in everyday language.
- Total cost of ownership (TCO)Total cost of ownership (TCO) is the full cost of a technology decision over its useful life, not just the purchase price or build quote.
05where insomnia club fits
Insomnia Club designs AI features with cost per task as a requirement: lean requests, the right model per step, caching where providers support it, and spend tracked from day one.
see custom AI development →tell us what keeps you up at night.
Scoped by the people who ship it. Priced before we start.
book a call drop your number