Introduction: The RAG Demo That Works Until It Meets Reality
Retrieval-Augmented Generation (RAG) has become one of the most practical approaches to building applications powered by Large Language Models (LLMs).
The idea is straightforward: retrieve relevant information from your own data, provide it to an LLM as context, and generate an answer grounded in that information.
Need a chatbot that answers questions about internal documentation? RAG can help.
Want an AI assistant that understands product manuals, company policies, API documentation, or support knowledge bases? RAG is a natural starting point.
The first implementation often looks deceptively simple:
- Split documents into chunks.
- Generate embeddings for each chunk.
- Store the embeddings in a vector database.
- Convert the user’s question into an embedding.
- Retrieve the nearest chunks.
- Send those chunks to the LLM and generate an answer.
You can build a working prototype with surprisingly little code.
But then production happens.
The chatbot retrieves a paragraph that contains the right keywords but misses the actual answer. A question about a specific error code returns a semantically similar but technically incorrect document. A policy question retrieves an outdated policy. Two users ask nearly identical questions and receive different answers because retrieval returns different context.
Increasing the number of retrieved chunks sometimes makes things worse. Switching to a more powerful LLM does not consistently solve the problem. Even changing the embedding model may produce only marginal improvements.
The problem is often not the LLM. It is the quality of the retrieval pipeline feeding it.
A vector database is an important component of a RAG system, but it is not the RAG system itself.
In this article, we’ll look at how to design retrieval systems that work beyond a successful demo.
1. Understanding the Real RAG Problem
Let’s use a practical example.
Imagine we are building an AI assistant for a ride-hailing platform. A user asks:
What happens if I cancel a ride after the driver has already arrived?
Our knowledge base contains:
- Ride cancellation policies
- Driver arrival and waiting-time rules
- Refund policies
- Customer support procedures
- Historical versions of those documents
A naïve RAG implementation might retrieve the five chunks most similar to the user’s question.
That sounds reasonable. But similarity does not guarantee correctness.
The system might retrieve a general cancellation policy while missing the specific section describing cancellation after driver arrival. It might retrieve an old policy that is no longer applicable. It might return a document intended for drivers when the question concerns riders.
The LLM can only work with the evidence it receives. If the correct policy never enters the context, even an excellent model may generate an incorrect answer.
This gives us a useful way to reason about RAG quality:
Answer quality = retrieval quality + context quality + generation quality
This is not a literal mathematical equation. It is an engineering model for understanding where failures originate.
If retrieval is poor, generation has to work with poor evidence. If retrieval is strong but context assembly is badly designed, relevant information may still be lost. If both are good, the LLM has a much better chance of generating a useful, grounded response.
The implication is important: we need to engineer the entire pipeline.
2. A Production-Grade RAG Architecture
Instead of thinking of RAG as a vector database connected to an LLM, think of it as a sequence of independently testable stages.

The ingestion pipeline
Before users ask questions, the system must prepare the knowledge base.
Step 1: Document processing
Extract useful text from PDFs, HTML pages, Markdown files, databases, and other sources. Preserve document structure whenever possible.
Step 2: Chunking
Split the content into meaningful units without destroying the relationships between sentences, sections, tables, or procedures.
Step 3: Metadata enrichment
Attach information such as document ID, title, section, source, version, tenant, access permissions, and publication date.
Step 4: Embedding generation
Convert each chunk into a vector representation using an embedding model.
Step 5: Indexing
Store the vectors, text, and metadata in a retrieval system such as Qdrant or PostgreSQL with pgvector.
The query-time pipeline
When a user asks a question, a different sequence of operations begins.
- Understand the query. Normalize terminology, resolve references where possible, or rewrite an ambiguous query.
- Retrieve candidates. Use semantic search, keyword search, or both.
- Apply metadata and access filters. Exclude documents the user should not see and documents that are out of scope.
- Rerank candidates. Use a stronger relevance model to identify the best evidence.
- Assemble context. Remove duplicates, preserve useful surrounding information, and stay within the model’s context budget.
- Generate the response. Ask the LLM to answer from the retrieved evidence and cite its sources.
- Evaluate and observe. Record retrieval quality, latency, token usage, failures, and feedback.
Not every application needs every stage. A small, well-structured knowledge base may work well with straightforward semantic search. A large enterprise knowledge base may require several retrieval strategies and more sophisticated ranking.
The architecture should follow the problem—not the other way around.
3. Chunking: The First Place Retrieval Quality Can Go Wrong
Chunking is one of the earliest design decisions in a RAG pipeline, and it has a direct effect on what the retrieval system can find.
Consider this paragraph:
“Customers can cancel a ride before the driver arrives. If the driver has already arrived, cancellation charges may apply according to the applicable waiting-time and cancellation policy.”
Suppose we split the paragraph into two chunks:
Chunk A
“Customers can cancel a ride before the driver arrives.”
Chunk B
“If the driver has already arrived, cancellation charges may apply according to the applicable waiting-time and cancellation policy.”
A question about cancellation after driver arrival may retrieve Chunk B, but it may miss the broader context that distinguishes the two situations.
This is a small example, but the same problem becomes harder with long technical documents, tables, API specifications, and policies containing cross-references.
Common chunking strategies

Fixed-size chunking
Split text into a fixed number of tokens or characters, optionally with overlap.
Advantages:
- Easy to implement
- Predictable chunk sizes
- Useful as a baseline
Disadvantages:
- Can break sentences, procedures, or tables
- Does not understand document structure
- May separate a rule from its conditions
Structure-aware chunking
Split according to headings, paragraphs, Markdown sections, HTML structure, or document boundaries.
Advantages:
- Preserves logical sections
- Often works well for technical documentation
- Makes citations and source navigation easier
Disadvantages:
- Section sizes can vary considerably
- Large sections may still need subdivision
Semantic chunking
Identify boundaries based on changes in meaning rather than only length.
Advantages:
- Can preserve conceptual coherence
- Useful for unstructured documents
Disadvantages:
- Adds processing complexity
- Depends on the segmentation method and content
Contextual chunking
Attach useful surrounding information to each chunk, such as the document title, section heading, or a short contextual description.
For example, instead of storing only:
“Cancellation charges may apply after arrival.”
Store something closer to:
“Document: Rider Cancellation Policy. Section: Cancellation after driver arrival. Rule: Cancellation charges may apply after the driver arrives, subject to the applicable waiting-time policy.”
The extra context can help retrieval distinguish a relevant policy from similar text elsewhere in the knowledge base.
Contextual enrichment should preserve the source’s meaning. Generated summaries or descriptions must not introduce new policy claims.
Is a 500-token chunk the right size?
There is no universal answer.
A 500-token chunk with 100-token overlap can be a reasonable baseline, but the right configuration depends on the content and the retrieval task.
For API documentation, keeping an endpoint’s description, parameters, and relevant constraints together may be more useful than enforcing a strict token count. For FAQs, each question-answer pair may be a natural retrieval unit. For long policy documents, heading-aware chunks may be more appropriate.
I would start with a simple, reproducible strategy, establish a baseline, and then compare alternatives using an evaluation dataset.
The right chunking strategy is the one that improves retrieval for your actual queries—not the one that looks best in a tutorial.
4. Semantic Search Is Powerful, but It Is Not Enough
Vector search retrieves content based on semantic similarity.
An embedding model converts a query and a document chunk into vectors. A similarity function then estimates how closely they relate.
This is useful because users do not always use the same words as the source documents.
For example:
- Query: “How do I get my money back?”
- Document: “Refund eligibility and payment reversal procedure”
The wording differs, but the concepts overlap.
However, semantic similarity can struggle with exact identifiers and domain-specific terminology.
Consider these queries:
- “What does error
E1027mean?” - “Show API endpoint
/v2/rides/{rideId}/cancel.” - “What is the retention period for policy
POL-481?” - “Does
X-Request-Idneed to be included?”
The exact identifier may be more important than the general meaning of the sentence. A semantically similar result is not necessarily the correct result.
This is where hybrid retrieval becomes useful.
5. Hybrid Search: Combine Meaning with Exact Matching
Hybrid retrieval combines two complementary techniques:
Dense vector search finds semantically related content.
Lexical search, commonly implemented with a BM25-based search engine, finds content based on matching terms and their importance.
Neither is universally better.
Vector search is useful when users paraphrase concepts. Lexical search is valuable when exact names, identifiers, product codes, or technical phrases matter.
A hybrid pipeline can execute both searches and combine their results.
For example, consider a question about error E1027. Vector search may retrieve general troubleshooting guidance, while lexical search finds the exact section documenting that error. Combining the results improves the chance of finding the correct evidence.
A practical hybrid retrieval flow
- Accept the user’s query.
- Run dense retrieval to find semantically related chunks.
- Run lexical retrieval to find exact or important term matches.
- Merge the candidate lists.
- Deduplicate by document or chunk identity.
- Apply relevance and access constraints.
- Rerank the candidates.
- Select the final context for the LLM.
A simple conceptual implementation in Java might look like this:
@Service
public class HybridSearchService {
private final VectorSearchClient vectorSearchClient;
private final KeywordSearchClient keywordSearchClient;
private final ResultMerger resultMerger;
private final Reranker reranker;
public HybridSearchService(
VectorSearchClient vectorSearchClient,
KeywordSearchClient keywordSearchClient,
ResultMerger resultMerger,
Reranker reranker) {
this.vectorSearchClient = vectorSearchClient;
this.keywordSearchClient = keywordSearchClient;
this.resultMerger = resultMerger;
this.reranker = reranker;
}
public List<SearchResult> search(
String query,
SearchFilters filters,
int topK) {
var semanticResults =
vectorSearchClient.search(query, filters, 30);
var keywordResults =
keywordSearchClient.search(query, filters, 30);
var candidates = resultMerger.merge(
semanticResults,
keywordResults
);
return reranker.rerank(query, candidates)
.stream()
.limit(topK)
.toList();
}
}
The interfaces above are application-level abstractions, not built-in Spring AI classes. The vector and keyword clients can be implemented using Qdrant, pgvector, Elasticsearch, OpenSearch, or another suitable backend.
For production implementations, the two retrieval calls can run concurrently, provided they share the same security and metadata constraints.
The 30 candidate limit is illustrative, not a universal default. It should be tuned against recall, latency, cost, and the reranker’s capacity.
How should the results be combined?
One option is to normalize scores and calculate a weighted combination:
combinedScore = α × vectorScore + β × keywordScore
This approach requires care because scores from different retrieval systems may not be directly comparable.
Another option is Reciprocal Rank Fusion (RRF). It combines ranked lists based on each result’s position rather than relying on raw score comparability.
A simplified RRF score is:
RRF(d) = Σ 1 / (k + rankᵢ(d))
Here, rankᵢ(d) is the position of document d in a result list, and k is a smoothing constant.
RRF is a useful baseline when combining rankings from different retrieval methods. It is not guaranteed to be optimal for every domain, but it avoids some of the problems caused by mixing uncalibrated scores.
One implementation detail matters: PostgreSQL full-text search does not provide BM25 as its standard ranking function. If BM25 is a requirement, choose an engine or extension that explicitly supports it, or use PostgreSQL’s built-in full-text ranking as a separate lexical retrieval strategy.
6. Metadata Filtering: Relevance Is Not the Only Requirement
Imagine an enterprise RAG platform serving several companies.
Two companies might have documents with almost identical content. A user from Company A asks a question, and the vector search returns a highly relevant document belonging to Company B.
From a similarity perspective, the result might be excellent.
From a security perspective, it is unacceptable.
Retrieval must respect authorization boundaries, tenant isolation, document status, and other application constraints.
Useful metadata can include:
- Tenant or organization ID
- Document ID and source
- Document version and publication status
- Access-control attributes
- Content type or product
- Language
- Effective date and expiry date
For example, the retrieval operation should conceptually apply a filter such as:
tenant_id = authenticated_user.tenant_id
AND status = "PUBLISHED"
AND effective_from <= current_time
AND (effective_to IS NULL OR effective_to > current_time)
This is illustrative filter logic; actual query syntax depends on the storage engine.
Security filters must be enforced by trusted server-side logic. Do not rely on the LLM to decide whether a user is allowed to see a document, and do not accept arbitrary tenant identifiers from an untrusted client as authorization.
Filtering can happen during candidate retrieval, and additional validation can happen before context assembly. Applying filters early is generally preferable because it prevents unauthorized documents from entering the candidate set in the first place.
There is also a performance trade-off. Highly selective filters can change the behavior and cost of approximate nearest-neighbor search, so filtering strategy should be tested with realistic data distributions.
7. Reranking: Finding the Best Evidence Among the Candidates

Hybrid retrieval improves candidate discovery, but the combined results may still contain weak matches.
A reranker adds a more detailed relevance assessment.
The first-stage retrievers are generally optimized for efficient candidate generation. A reranker evaluates a smaller set of candidates against the original query, often using a cross-encoder or another relevance model.
The process looks like this:
- Retrieve 30–100 candidate chunks using one or more retrieval methods.
- Score the candidates against the user’s query.
- Sort the candidates by relevance.
- Select the strongest evidence for context assembly.
The candidate counts are examples, not prescriptions. Larger sets can improve the chance of finding relevant material, but they also increase latency and reranking cost.
Why reranking helps
Suppose a query is:
“Can a rider cancel after the driver reaches the pickup location without paying a fee?”
The initial search might return:
- A general ride cancellation policy
- A driver cancellation policy
- A refund FAQ
- A policy about cancellation after driver arrival
- A document about scheduled rides
All five documents are related to cancellation, but only one may directly answer the question.
A reranker can help promote that document because it evaluates the relationship between the complete query and each candidate, rather than relying only on first-stage retrieval scores.
Reranking is not magic. It cannot retrieve a document that the candidate-generation stage failed to find, and it can still rank incorrect evidence highly. That is why retrieval recall and reranking quality should be measured separately.
8. Query Understanding and Context Assembly

Users rarely ask questions in the exact form used by the knowledge base.
Consider a conversation:
- “What is the cancellation policy?”
- “What if the driver is already there?”
- “Does that apply to scheduled rides too?”
The second and third questions depend on conversation history. Searching only the latest message may produce weak results because “that” and “there” are ambiguous.
A query-understanding stage can rewrite the latest question into a self-contained search query, using relevant conversation context.
For example:
“Does the cancellation policy apply to scheduled rides when the driver has arrived?”
Query rewriting should be used selectively. For exact identifiers or short technical queries, unnecessary rewriting can remove important terms or introduce ambiguity. In those cases, retaining the original query alongside a rewritten version may be useful.
After retrieval and reranking, the system still needs to assemble the context carefully.
More context is not always better.
A large prompt containing many marginally relevant chunks can bury the most useful evidence, increase token cost, and introduce conflicting information.
A good context assembler should:
- Remove duplicate or near-duplicate chunks.
- Preserve useful section headings and source references.
- Prefer authoritative and current documents.
- Keep related conditions and exceptions together.
- Respect the model’s context budget.
- Retain enough provenance to cite the source of each claim.
The objective is not to pass the maximum amount of information to the model. It is to pass the most useful evidence for the current question.
9. How Do We Know Retrieval Is Actually Improving?
This is where many RAG implementations are weaker than they appear.

A few manually tested questions can demonstrate that the application works, but they do not establish that it works consistently.
Before changing chunk size, switching embedding models, or adding reranking, build an evaluation dataset.
For each test query, record:
- The question
- The expected relevant document or chunk IDs
- The expected answer, where appropriate
- The source or policy version
- The relevant tenant or access constraints
Then compare retrieval configurations against the same dataset.
Metrics worth measuring
Recall@K
Measures whether the relevant items appear in the top K results.
If the correct policy appears somewhere in the top 10 results, Recall@10 counts that as a successful retrieval for that relevant item.
Mean Reciprocal Rank (MRR)
Measures how highly the first relevant result appears. A correct result at rank 1 receives more credit than one at rank 10.
Normalized Discounted Cumulative Gain (nDCG)
Measures ranking quality when results have graded relevance, giving greater weight to highly relevant results near the top.
Answer groundedness
Evaluates whether the generated answer is supported by the retrieved evidence. This requires its own evaluation methodology; a good retrieval score alone does not guarantee a grounded answer.
Operational metrics
Track p50 and p95 latency, token usage, model cost, error rates, empty retrievals, and user feedback.
The exact metrics depend on the product. A support assistant may prioritize grounded answers and escalation rates. A developer documentation assistant may prioritize exact endpoint and error-code retrieval. A policy assistant may require strict version correctness and access-control compliance.
A simple experiment matrix
Rather than changing several variables at once, compare configurations systematically.
| Experiment | Retrieval strategy | What it tells us |
|---|---|---|
| A | Dense vector search | Establishes a baseline |
| B | Improved chunking + dense search | Tests chunking quality |
| C | Hybrid retrieval | Measures the value of lexical search |
| D | Hybrid retrieval + reranking | Tests ranking improvements |
| E | Hybrid retrieval + reranking + query rewriting | Measures the additional value of query understanding |
Use the same evaluation dataset and report both quality and operational impact.
For example, a configuration might improve Recall@10 but increase p95 latency and cost. Whether that is acceptable depends on the product’s requirements.
The point is not to maximize one metric at any cost. It is to make retrieval trade-offs visible and measurable.
10. A Practical Implementation Path with Java and Spring Boot
For engineers building RAG applications in the Java ecosystem, I would implement the system in small, testable stages.
Stage 1: Build a baseline
Use Spring Boot, an embedding model, and Qdrant or PostgreSQL with pgvector. Implement document ingestion, metadata storage, vector retrieval, and source-aware answers.
Stage 2: Improve chunking
Preserve headings and document structure. Add chunk identifiers, source references, and document version metadata.
Stage 3: Add hybrid retrieval
Introduce a lexical search backend or capability and merge its candidates with vector results.
Stage 4: Add reranking
Evaluate a reranker on the candidate set. Measure the relevance gain and additional latency.
Stage 5: Add evaluation and observability
Build a repeatable test suite and capture retrieval traces, selected source IDs, latency, token consumption, and failure categories.
Stage 6: Harden for production
Enforce access control, document freshness, deletion and re-indexing workflows, timeouts, bounded retries, and sensible fallback behavior.
Keep the responsibilities separate:
DocumentIngestionServiceowns ingestion and chunking.EmbeddingServiceabstracts embedding generation.VectorSearchClientowns dense retrieval.KeywordSearchClientowns lexical retrieval.HybridSearchServicecombines candidates.Rerankerorders the candidates.ContextAssemblerprepares evidence for the LLM.RagEvaluationServicemeasures retrieval and answer quality.
These are example application boundaries, not mandatory class names. The important principle is to make each stage independently replaceable and testable.
For a small application, this may seem like more structure than necessary. For a production system where retrieval quality affects user trust, debugging, and operating cost, these boundaries make the system much easier to evolve.
Conclusion: Engineer the Retrieval Pipeline, Not Just the Vector Store
RAG is often introduced as an embeddings-and-vector-search problem because that is the easiest part to demonstrate.
Production systems are different.
The quality of the final answer depends on whether the right evidence can be found, whether it is ranked correctly, whether access rules are respected, and whether the context is assembled in a way the LLM can use.
A vector database is a useful building block. It is not a substitute for retrieval engineering.
If I were starting a RAG implementation today, I would focus on five things before reaching for a more sophisticated model:
- Preserve document meaning through sensible chunking.
- Combine semantic retrieval with lexical search when the domain requires it.
- Rerank candidates when first-stage ranking is not good enough.
- Apply authorization and metadata constraints throughout retrieval.
- Build an evaluation dataset before attempting to optimize quality.
Most importantly, I would treat retrieval as an independently measurable subsystem rather than a hidden step inside a prompt.
That shift—from wiring components together to engineering and measuring the retrieval pipeline—is what turns a promising RAG demo into a system that can be trusted in production.
Next in this series: Stop Building Uncontrolled AI Agents: Apply Distributed Systems Principles to Agentic AI.