A language model may not know what is in a document it has never seen. Retrieval-augmented generation (RAG) addresses this by finding relevant information in an external source and giving it to the model along with the question before the model writes an answer.

How does RAG work?

  1. Documents are prepared and divided into searchable passages.
  2. When someone asks a question, a retrieval system finds relevant passages. It may use keyword search, numerical embeddings, or both.
  3. The question and retrieved passages are sent to a language model, which generates a response based on that context.

The model is not necessarily retrained whenever documents change. Instead, the application supplies information at answer time. For example, an internal assistant might search a company handbook before answering an employee’s question.

When is RAG useful?

It is useful when information is specialized, changes over time, or should be traceable to a document. Retrieval can reduce unsupported answers, but it does not guarantee accuracy. The system may retrieve an outdated or irrelevant passage, and the model may still misread it. Readers should check the cited source and its date.

Search normally returns results for a person to inspect. RAG adds a generation step: a model combines retrieved material into an answer. The quality of that answer depends on the quality of both retrieval and generation.

In short

RAG combines external information retrieval with language generation. Related concepts include embeddings, context windows, and AI hallucination.