AI Engineering#RAG#Vector Search#Embeddings#AI

RAG Didn't Solve My Hallucination Problem — Here's Why

Vector search gives you relevant chunks, not factual correctness. An empirical breakdown of chunking strategies, hybrid reranking, and self-corrective verification.

Arya
Arya
Full-Stack Product Builder & Engineer
Published on
•
9 min read
RAG Didn't Solve My Hallucination Problem — Here's Why
Share this dispatch:

The universal playbook for eliminating LLM hallucinations has been repeated at every AI conference for two years:

"Just chunk your documents, embed them into a vector database, retrieve top-k chunks, and inject them into the prompt. Problem solved."

Except it isn't. When we stress-tested a standard RAG pipeline against complex domain knowledge bases, hallucination rates remained stubborn at 18–24%. The model still confidently invented facts, misinterpreted ambiguous paragraphs, and synthesized contradictions from adjacent chunks.

Here is what went wrong, why cosine similarity alone is insufficient, and what actually fixed our accuracy.


1. The Cosine Similarity Trap

Vector embeddings are semantic approximations. A chunk that is semantically similar to a question is frequently not factually probative of the answer.

Consider this query:

"Does our plan cover water damage if the pipe burst occurred during winter freezing?"

A standard vector query often pulls chunks containing:

  • General policy limits on water damage.
  • Exclusions for seasonal freezing in vacant properties.
  • Normal plumbing maintenance guidelines.

Because all these passages share high lexical and semantic overlap with "water damage" and "freezing," cosine similarity scores them above 0.85. But none of them contain the critical exclusion clause needed to answer the question accurately.

python
# Naive Vector Retrieval
results = vector_db.similarity_search(
    query="Does our plan cover pipe bursts during winter?",
    k=4
)
# Returns semantically related fluff that confuses the context window

2. Moving to Hybrid Search with BM25 & Cohere Rerank

To fix precision, we introduced Hybrid Search combining dense vector representations with sparse keyword indices (BM25), followed by a cross-encoder reranker.

text
flowchart TD
    Q[User Question] --> Dense[Dense Vector Search]
    Q --> Sparse[BM25 Keyword Search]
    Dense --> Merge[Reciprocal Rank Fusion]
    Sparse --> Merge
    Merge --> Top50[Top 50 Candidates]
    Top50 --> Rerank[Cross-Encoder Reranker]
    Rerank --> Top5[Top 5 High-Precision Chunks]
    Top5 --> LLM[LLM Generation Node]

Why Cross-Encoders Matter

Bi-encoders (standard embeddings) generate vector representations of documents in isolation. Cross-encoders examine the question and passage simultaneously, computing attention across both texts.

The result is a dramatic leap in ranking precision:

TechniquePrecision@5Hallucination RateLatency
Pure Vector Search42.1%21.4%~40ms
Hybrid (BM25 + Dense)67.8%13.2%~75ms
Hybrid + Cross-Encoder Rerank89.5%3.8%~180ms

3. Chunking with Semantic Boundaries

Fixed-size chunking (e.g. 500 characters with 50 character overlap) regularly splits crucial sentences or cuts off dependent clauses halfway through.

Instead of naive token counting:

  • Chunk by markdown structural headers (H2, H3).
  • Preserve table rows together rather than splitting across chunks.
  • Append parent metadata (Document Title > Section Name) to each child chunk header.
typescript
interface EnrichedChunk {
  id: string;
  parentDocument: string;
  hierarchyPath: string[]; // ["Architecture", "Storage", "Failover"]
  content: string;
  tokens: number;
}

4. Self-Correction & Groundedness Checks

Before streaming the generated answer back to the client, a lightweight verification prompt checks whether every asserted factual claim in the response is directly grounded in the retrieved chunks.

If an assertion cannot be cited directly back to the source text, the generation is rejected and the system falls back to asking for clarification.


Conclusion

RAG is not a magical plug-and-play solution. Real reliability requires clean semantic chunking, hybrid keyword + vector retrieval, cross-encoder reranking, and rigorous verification steps. When implemented with discipline, hallucinations drop to near-zero.

Share this dispatch:
Further Reading

Related Dispatches

View all stories →