Retrieval-augmented generation (RAG)
Retrieval-augmented generation (RAG) is a technique in which an AI system first searches your own documents or data for passages relevant to a question, then gives those passages to a language model to write its answer. The model answers from your current, approved information instead of relying only on what it learned in training, and can cite sources.
01why it matters for a business
RAG is how most companies make a general model useful on their own knowledge: policies, contracts, product documentation, support history, standard operating procedures. It keeps answers current, because updating a document updates the answers without retraining anything. It allows citations, so users can check the source. And when built properly it respects access controls, because retrieval can be limited to what each user is allowed to see.
Most RAG failures are retrieval failures, not model failures. If the search step returns the wrong passages, the model will confidently answer from the wrong material. Quality depends on unglamorous work: cleaning documents, splitting them sensibly, keeping metadata such as dates and owners, combining semantic and keyword search, and testing with real questions.
02what it looks like in practice
A commercial HVAC contractor has years of equipment manuals, service bulletins, and technician notes. Field technicians ask an assistant on their phones how to clear a specific fault code on a specific unit. The system retrieves the matching manual section and recent bulletins for that model, and the model writes a short answer with links to the pages it used. When a bulletin is superseded, the old one is retired from the index and the answers change the same day.
03common mistakes
- Indexing everything, including outdated and conflicting versions, then wondering why answers contradict each other.
- Ignoring permissions. Retrieval must respect who is allowed to see which document.
- No citations. Users should be able to see and check the source of every answer.
- Testing with easy questions. Build a test set from real questions, including ones the documents cannot answer, and check that the system says so.
04related terms
- EmbeddingsAn embedding is a list of numbers that represents the meaning of a piece of text, an image, or other data, produced by an embedding model.
- Vector databaseA vector database is a database designed to store embeddings and quickly find the ones most similar to a query, a task called similarity or nearest-neighbor search.
- AI hallucinationA hallucination is output from an AI model that is fluent and confident but false or unsupported: an invented fact, a made-up citation, a policy that does not exist, or a number that appears nowhere in the source.
- Fine-tuningFine-tuning is the process of further training an existing AI model on a curated set of your own examples so it consistently produces a particular style, format, or behavior.
- AI knowledge baseAn AI knowledge base is a company's documents, policies, procedures, and records organized so an AI assistant can search them and answer questions with citations.
05where insomnia club fits
Retrieval over your own documents and records is a core part of Insomnia Club's AI work: the data layer, permissions, citations, and the evaluation set that proves answers are right before launch.
see custom AI development →tell us what keeps you up at night.
Scoped by the people who ship it. Priced before we start.
book a call drop your number