RAG Systems

Document Preparation for RAG Systems

Master the art of preparing documents for Retrieval-Augmented Generation. Learn chunking strategies, embedding optimization, and retrieval best practices.

Markdown Converters team
January 10, 2025
14 min read
⚠️ AI-Generated Content Notice

This article was generated with AI assistance. While we strive for accuracy, please verify any technical implementations or system configurations independently before deploying to production.

Retrieval-Augmented Generation (RAG) systems combine the power of large language models with your specific knowledge base. But RAG is only as good as your document preparation. This guide covers everything you need to know.

What is RAG and Why Document Prep Matters

RAG in 60 Seconds

RAG systems work in three steps:

  1. Retrieval: Find relevant documents from your knowledge base
  2. Augmentation: Add retrieved context to the user's query
  3. Generation: LLM generates response using both query and context

Why Preparation is Critical

Poor document preparation leads to:

  • Irrelevant retrievals: System finds wrong information
  • Context loss: Important details get separated
  • Slow performance: Inefficient chunking slows retrieval
  • Inaccurate responses: LLM works with incomplete context
  • High costs: Wasted tokens and compute resources

Good preparation = Better retrieval = More accurate responses

Step 1: Convert to Clean Markdown

Why Markdown for RAG?

  • Semantic structure: Headings help identify topic boundaries
  • Token efficiency: Minimal syntax overhead
  • Easy parsing: Simple to chunk and process
  • Metadata preservation: Structure aids retrieval
  • Human readable: Easy to debug and maintain

Conversion Best Practices

  1. Start with quality sources: Use our ePub converter or PDF converter
  2. Preserve structure: Maintain heading hierarchy
  3. Clean artifacts: Remove page numbers, headers, footers
  4. Fix formatting: Ensure consistent markdown syntax
  5. Validate output: Check for conversion errors

Step 2: Add Metadata and Frontmatter

Metadata dramatically improves retrieval accuracy:

---
document_id: atomic_habits_ch1
title: The Surprising Power of Atomic Habits
source: Atomic Habits by James Clear
chapter: 1
category: personal_development
topics:
  - habit formation
  - behavior change
  - compound effect
keywords:
  - atomic habits
  - 1% improvement
  - identity change
  - habit loop
date_added: 2024-01-15
content_type: book_chapter
reading_level: intermediate
estimated_tokens: 2500
---

# The Surprising Power of Atomic Habits

[Content begins...]

Essential Metadata Fields

  • document_id: Unique identifier
  • title: Clear, descriptive title
  • source: Original source attribution
  • category: Primary topic area
  • topics: Specific subjects covered
  • keywords: Search terms and concepts
  • content_type: Book, article, guide, etc.

Optional but Useful

  • date_added: When document was added
  • last_updated: Most recent modification
  • author: Original author
  • reading_level: Complexity indicator
  • estimated_tokens: Size for context planning
  • related_docs: Links to related content

Step 3: Chunking Strategy

Chunking is the most critical step in RAG document preparation. Get this wrong and everything else fails.

What is Chunking?

Breaking large documents into smaller pieces (chunks) that can be:

  • Embedded as vectors
  • Retrieved independently
  • Fit within context windows
  • Provide complete, coherent information

Chunking Strategies Compared

1. Fixed-Size Chunking

Method: Split every N tokens/characters

Chunk 1: [0-500 tokens]
Chunk 2: [500-1000 tokens]
Chunk 3: [1000-1500 tokens]

Pros:

  • Simple to implement
  • Predictable chunk sizes
  • Fast processing

Cons:

  • Breaks mid-sentence or mid-concept
  • Loses semantic coherence
  • Poor retrieval accuracy

Best for: Quick prototypes only

2. Semantic Chunking (Recommended)

Method: Split at natural boundaries (headings, paragraphs, topics)

Chunk 1: ## Introduction + full section
Chunk 2: ## Key Concept 1 + full section
Chunk 3: ## Key Concept 2 + full section

Pros:

  • Preserves semantic meaning
  • Natural topic boundaries
  • Better retrieval accuracy
  • More coherent context

Cons:

  • Variable chunk sizes
  • Requires document structure
  • More complex implementation

Best for: Most RAG applications

3. Sliding Window Chunking

Method: Overlapping chunks with shared context

Chunk 1: [0-500 tokens]
Chunk 2: [400-900 tokens] (100 token overlap)
Chunk 3: [800-1300 tokens] (100 token overlap)

Pros:

  • Reduces context loss at boundaries
  • Improves retrieval recall
  • Handles concepts spanning boundaries

Cons:

  • Increases storage requirements
  • More embeddings to compute
  • Potential redundancy in results

Best for: Critical applications where accuracy matters most

4. Hierarchical Chunking

Method: Multiple levels of granularity

Level 1: Entire document summary
Level 2: Chapter summaries
Level 3: Section content
Level 4: Paragraph details

Pros:

  • Retrieves at appropriate level
  • Can provide broad or specific context
  • Excellent for large documents

Cons:

  • Complex to implement
  • Requires careful hierarchy design
  • More storage and processing

Best for: Large knowledge bases with complex structure

Recommended Chunk Sizes

Short-form Q&A200-400 tokens

Quick facts, definitions, simple answers

General purpose (Recommended)400-800 tokens

Most use cases, balanced context and specificity

Long-form content800-1500 tokens

Complex topics, detailed explanations

Avoid>2000 tokens

Too large, reduces retrieval precision

Chunking Implementation Example

Semantic chunking by heading level:

# Document Title (not chunked separately)

## Chapter 1: Introduction
[Full section content = Chunk 1]

### Subsection 1.1
[If chapter too long, subsection = Chunk 2]

### Subsection 1.2
[Subsection = Chunk 3]

## Chapter 2: Main Content
[Full section = Chunk 4]

Step 4: Enhance Chunks with Context

Problem: Context Loss

When you chunk, you lose surrounding context:

Original: 
# Atomic Habits
## Chapter 1
### The Four Laws
1. Make it obvious

Chunk becomes just:
"1. Make it obvious"
(Missing: book title, chapter, section)

Solution: Context Enrichment

Add hierarchical context to each chunk:

---
document: Atomic Habits by James Clear
chapter: 1 - The Surprising Power of Atomic Habits
section: The Four Laws of Behavior Change
subsection: Law 1 - Make It Obvious
---

# Law 1: Make It Obvious

The first law of behavior change is to make it obvious...

[Full subsection content]

---
Related concepts: habit loop, cue, environment design
Previous: Introduction to Four Laws
Next: Law 2 - Make It Attractive

Context Enrichment Strategies

  1. Hierarchical breadcrumbs: Book → Chapter → Section
  2. Related concepts: Link to related chunks
  3. Navigation hints: Previous/next sections
  4. Summary: Brief section summary
  5. Keywords: Important terms in this chunk

Step 5: Generate Embeddings

What are Embeddings?

Vector representations of text that capture semantic meaning. Similar concepts have similar vectors.

Embedding Model Selection

OpenAI text-embedding-3-large

3072 dimensions, excellent quality

Best for: High-accuracy applications, budget available

OpenAI text-embedding-3-small

1536 dimensions, good balance

Best for: Most applications, cost-conscious

Sentence Transformers (open-source)

384-768 dimensions, free

Best for: Self-hosted, privacy-sensitive

Embedding Best Practices

  • Consistent model: Use same model for all documents
  • Include metadata: Embed title + content together
  • Batch processing: Generate embeddings in batches for efficiency
  • Cache embeddings: Store vectors, don't regenerate
  • Version control: Track which embedding model version used

Step 6: Optimize for Retrieval

Retrieval Strategies

1. Semantic Search (Vector Similarity)

Find chunks with similar meaning to query:

  • Cosine similarity between query and chunk embeddings
  • Return top K most similar chunks
  • Works well for conceptual queries

2. Keyword Search (BM25)

Traditional keyword matching:

  • Exact term matching with relevance scoring
  • Works well for specific terms or names
  • Fast and efficient

3. Hybrid Search (Recommended)

Combine semantic and keyword search:

Results = (0.7 × Semantic Score) + (0.3 × Keyword Score)
  • Best of both worlds
  • Handles both conceptual and specific queries
  • More robust retrieval

Retrieval Optimization Techniques

  1. Query expansion: Rephrase query multiple ways
  2. Reranking: Use LLM to rerank initial results
  3. Metadata filtering: Pre-filter by category, date, etc.
  4. MMR (Maximal Marginal Relevance): Diversify results
  5. Parent-child retrieval: Retrieve small chunks, return larger context

Step 7: Test and Iterate

Create Test Queries

Build a test set of queries and expected results:

Test Cases:
1. Query: "How do I build a new habit?"
   Expected: Chunks from Atomic Habits, Four Laws section
   
2. Query: "What's the two-minute rule?"
   Expected: Specific section on two-minute rule
   
3. Query: "Compare habit stacking and implementation intentions"
   Expected: Multiple chunks from different sections

Evaluation Metrics

  • Precision@K: How many of top K results are relevant?
  • Recall: Did we retrieve all relevant chunks?
  • MRR (Mean Reciprocal Rank): Position of first relevant result
  • NDCG: Normalized discounted cumulative gain
  • Human evaluation: Manual review of response quality

Common Issues and Fixes

Problem: Irrelevant Results

System retrieves wrong chunks

Solutions:

  • • Improve chunk boundaries
  • • Add more metadata
  • • Use hybrid search
  • • Implement reranking

Problem: Incomplete Context

Retrieved chunks lack necessary info

Solutions:

  • • Increase chunk size
  • • Add overlapping chunks
  • • Include hierarchical context
  • • Retrieve parent chunks

Advanced Techniques

1. Multi-Vector Retrieval

Generate multiple embeddings per chunk:

  • One for main content
  • One for summary
  • One for questions it answers

2. Hypothetical Document Embeddings (HyDE)

Generate hypothetical answer to query, embed that, then search:

  1. User asks: "How do I build a habit?"
  2. LLM generates hypothetical answer
  3. Embed hypothetical answer
  4. Search for similar real chunks

3. Query Decomposition

Break complex queries into sub-queries:

Complex Query: "Compare habit stacking and implementation intentions, 
and explain when to use each"

Decomposed:
1. "What is habit stacking?"
2. "What are implementation intentions?"
3. "When should I use habit stacking?"
4. "When should I use implementation intentions?"

Tools and Frameworks

Vector Databases

  • Pinecone: Managed, easy to use, scales well
  • Weaviate: Open-source, feature-rich
  • Qdrant: Fast, good for production
  • ChromaDB: Simple, good for prototyping
  • FAISS: Facebook's library, self-hosted

RAG Frameworks

  • LangChain: Most popular, extensive features
  • LlamaIndex: Specialized for RAG, excellent docs
  • Haystack: Production-ready, modular
  • txtai: Lightweight, easy to start

Production Checklist

Before deploying your RAG system:

  • ✓ Documents converted to clean markdown
  • ✓ Comprehensive metadata added
  • ✓ Semantic chunking implemented
  • ✓ Context enrichment applied
  • ✓ Embeddings generated and stored
  • ✓ Hybrid search configured
  • ✓ Test queries passing
  • ✓ Evaluation metrics tracked
  • ✓ Monitoring and logging set up
  • ✓ Update process documented

Conclusion

Document preparation is the foundation of successful RAG systems. Invest time in proper conversion, chunking, and enrichment—it directly impacts your system's accuracy and usefulness.

Start with semantic chunking at 400-800 tokens, add comprehensive metadata, and test thoroughly. Iterate based on real-world performance.

Prepare Documents for Your RAG System

Convert your documents to clean, structured markdown optimized for RAG retrieval.