Ask a general-purpose language model about your company’s current pricing tiers or last week’s product changes, and it will either say it doesn’t know or, worse, confidently make something up. The model was trained on a fixed snapshot of data, so anything that happened after that snapshot, or anything private to your organization, is invisible to it. Retrieval-augmented generation (RAG) fixes this by retrieving relevant external data before the model generates a response, grounding the answer in real, current content instead of only what the model memorized during training.

Understanding Retrieval Augmented Generation (RAG)

The Basics of RAG

Retrieval Augmented Generation (RAG) is a framework that integrates two distinct yet complementary AI methodologies: retrieval-based models and generative models.

  1. Retrieval-Based Models: These models excel in fetching relevant information from a large corpus of data. Given a query, a retrieval-based model searches through a database to find documents or passages that are most pertinent to the query.
  2. Generative Models: These models, often based on transformer architectures like GPT-3, are designed to generate coherent and contextually appropriate text. They are capable of producing novel content based on the input they receive.

RAG combines these two approaches by first retrieving relevant documents or passages from a large dataset and then using this retrieved information to inform and enhance the generation of responses.

How RAG Works

The RAG framework typically involves the following steps:

  1. Query Processing: When a query is received, the retrieval component of the RAG model searches a pre-indexed database to find the most relevant documents or passages.
  2. Contextual Embedding: The retrieved documents are then converted into embeddings, which are vector representations that capture the semantic meaning of the text.
  3. Response Generation: The generative model combines the original query with the embeddings of the retrieved documents to generate a response grounded in that retrieved content.

Where RAG Pipelines Run Into a Routing Problem

A RAG pipeline that only searches unstructured documents works fine with a single vector store. Real applications rarely stay that simple. A support bot that needs to answer “What’s the status of order #48213 and does our return policy cover it?” has to pull a structured order record from one system and an unstructured policy document from another, then combine both in a single response.

Most teams solve this by running a vector database for embeddings alongside a separate relational database for structured records, then writing application code to query both and merge the results. Every new data source becomes a new system to provision, secure, and keep in sync.

How TiDB Removes the Store-Per-Data-Type Decision

티DB stores structured relational data and vector embeddings in the same MySQL-compatible database, using a native VECTOR column type alongside ordinary SQL columns. A single query can filter on a structured field and rank by vector similarity in one statement:

SELECT order_id, status, description
FROM support_docs
WHERE order_id = 48213
ORDER BY VEC_Cosine_Distance(description_vec, :query_embedding)
LIMIT 5;

Because both the structured order record and the unstructured policy text live in the same schema, the routing decision from the previous section disappears. This is the same vector and transactional infrastructure that runs in production behind TiDB Cloud 스타터-based AI agent platforms like Kimi’s agent hosting service, which relies on the same underlying TiDB Cloud architecture at much larger scale than a typical RAG pipeline needs.

자주 묻는 질문

What is the difference between RAG and fine-tuning?

  • Fine-tuning updates the model’s weights on new data, which is expensive and needs retraining whenever the data changes.
  • RAG keeps the model unchanged and instead retrieves current data at query time.
  • RAG is typically cheaper to keep up to date, since updating a database is far less work than retraining a model.

Do I need a dedicated vector database to build RAG?

  • No. A relational database with a native vector column, like TiDB, can store embeddings alongside structured data.
  • This avoids running and syncing a separate vector database for a large class of RAG use cases.
  • A dedicated vector database still makes sense for pipelines that only ever need pure vector search at very large scale.

Does RAG eliminate hallucinations completely?

  • No. RAG substantially reduces hallucination by grounding responses in retrieved data, but it does not guarantee zero errors.
  • The model can still misinterpret or poorly summarize what was retrieved.
  • Retrieval quality matters as much as the model itself; irrelevant or outdated retrieved documents produce weak answers regardless of the generative model’s quality.

What is the difference between RAG and GraphRAG?

  • RAG retrieves relevant text chunks based on vector similarity.
  • GraphRAG additionally indexes entities and relationships as a knowledge graph, so it can answer multi-hop questions RAG alone cannot.
  • See the GraphRAG guide for the knowledge-graph-specific version of this pattern.

Get Started with RAG on TiDB

RAG works by grounding a model’s responses in data it can retrieve at query time, not just what it memorized during training. Storing that retrieved data, structured and vector alike, in one MySQL-compatible database removes the separate-store routing problem most RAG tutorials never address. TiDB Cloud 스타터, TiDB’s free tier, includes native vector search alongside standard SQL. Explore the full vector search and RAG pipelines documentation, or see the Chat with URL demo for a working example using LlamaIndex.


Last updated 9월 30, 2026

Spin up a database with 25 GiB free resources.

Start Right Away

💬 Let’s Build Better Experiences — Together

Join our Discord to ask questions, share wins, and shape what’s next.

Join Now