Fine-tuning
Fine-tuning is the process of further training an existing AI model on a curated set of your own examples so it consistently produces a particular style, format, or behavior. It changes the model's weights, unlike prompting or retrieval, which only change what the model is given at the moment it answers. It suits narrow, repeated tasks.
01why it matters for a business
Fine-tuning is often proposed as the way to teach a model about your business. It is usually the wrong tool for that. Facts that change, such as prices, policies, and inventory, belong in retrieval, where updates take effect immediately and answers can cite sources. Fine-tuning shines at behavior: a consistent output format, a specific tone, a specialized classification, or getting a smaller, cheaper model to match a larger one on a narrow task.
It also carries ongoing costs. You need a labeled dataset of good examples, an evaluation set to prove the tuned model is better, and a plan to repeat the process when the base model is updated or your requirements change. For many companies, better instructions and examples in the prompt deliver most of the benefit with none of the upkeep, so try those first.
02what it looks like in practice
A freight broker classifies thousands of carrier emails a day into a few dozen categories. A large model does it well but is slow and costly at that volume. The team collects a few thousand emails already labeled by staff, fine-tunes a small model, and compares both on a held-out test set. The tuned small model matches the large one on this task at a fraction of the cost per email, and anything it classifies with low confidence still goes to the larger model.
03common mistakes
- Fine-tuning to inject facts that change. Use retrieval for knowledge and fine-tuning for behavior.
- Training on messy or inconsistent examples. The model learns the inconsistency.
- No held-out test set, so nobody can show the tuned model is actually better.
- Forgetting maintenance. Base models and requirements change, and the tuning has to be redone.
04related terms
- Retrieval-augmented generation (RAG)Retrieval-augmented generation (RAG) is a technique in which an AI system first searches your own documents or data for passages relevant to a question, then gives those passages to a language model to write its answer.
- Prompt engineeringPrompt engineering is the practice of writing and refining the instructions, context, and examples given to an AI model so it produces reliable, useful output for a specific task.
- AI evaluations (evals)Evals are structured tests that measure how well an AI system performs on a defined task, using a set of real or realistic inputs with known good outcomes.
- Open-weight modelAn open-weight model is an AI model whose trained parameters, called weights, are published so anyone can download and run it on their own hardware or cloud, subject to its license.
05where insomnia club fits
Insomnia Club will tell you whether fine-tuning is worth it for your task and, when it is, builds the dataset, the evaluation, and the retraining process along with the model.
see custom AI development →tell us what keeps you up at night.
Scoped by the people who ship it. Priced before we start.
book a call drop your number