Large language model (LLM)
A large language model (LLM) is an AI model trained on very large amounts of text to predict the next piece of text, which lets it write, summarize, translate, classify, extract information, and reason through problems in everyday language. Claude, GPT, Gemini, and Llama are families of LLMs, the engine inside most modern AI products.
01why it matters for a business
For most business uses of AI today, an LLM is the core component, and the practical consequences follow from how it works. It generates the most plausible continuation, not a verified fact, so it can be fluent and wrong. It knows nothing about your company unless you give it your data at the moment it answers. And it is billed by the Token, so cost scales with how much text goes in and out.
Executives do not need the mathematics, but they do need a sound mental model: an LLM is a capable generalist with no memory of your business and no built-in way to check its own facts. Good systems add the missing pieces: retrieval of your data, tools to act, tests to catch errors, and people at the right checkpoints.
02what it looks like in practice
A logistics company wants to cut the time account managers spend writing weekly customer updates. An LLM is given the shipment data for each account, a template of what customers care about, and examples of good past updates. It drafts each update in seconds; the account manager edits and sends. The model does what LLMs are best at, turning structured facts into clear prose, while the facts themselves come from the company's systems, not the model's memory.
03common mistakes
- Asking an LLM for facts it was never given, such as your pricing or policies, and trusting the answer.
- Picking a model on benchmark headlines. Test candidates on your own tasks; rankings shift often and may not reflect your work.
- Ignoring data terms. Check whether a provider retains or trains on your inputs under the plan you actually use.
- Assuming bigger is always better. Smaller, cheaper models handle many routine tasks well.
04related terms
- TokenA token is the unit of text a language model reads and writes: a whole word, part of a word, a number, or a punctuation mark.
- Context windowA context window is the maximum amount of text, measured in tokens, that a language model can take into account in a single request, including your instructions, any documents or conversation history, and the answer it writes.
- AI hallucinationA hallucination is output from an AI model that is fluent and confident but false or unsupported: an invented fact, a made-up citation, a policy that does not exist, or a number that appears nowhere in the source.
- Retrieval-augmented generation (RAG)Retrieval-augmented generation (RAG) is a technique in which an AI system first searches your own documents or data for passages relevant to a question, then gives those passages to a language model to write its answer.
- Open-weight modelAn open-weight model is an AI model whose trained parameters, called weights, are published so anyone can download and run it on their own hardware or cloud, subject to its license.
05where insomnia club fits
Insomnia Club builds custom software around LLMs and chooses the model per task, testing candidates on your real work so the choice rests on evidence rather than leaderboards.
see custom AI development →tell us what keeps you up at night.
Scoped by the people who ship it. Priced before we start.
book a call drop your number