This article was generated with AI assistance. While we strive for accuracy, please verify any technical implementations or system configurations independently before deploying to production.
Retrieval-Augmented Generation (RAG) systems combine the power of large language models with your specific knowledge base. But RAG is only as good as your document preparation. This guide covers everything you need to know.
What is RAG and Why Document Prep Matters
RAG in 60 Seconds
RAG systems work in three steps:
- Retrieval: Find relevant documents from your knowledge base
- Augmentation: Add retrieved context to the user's query
- Generation: LLM generates response using both query and context
Why Preparation is Critical
Poor document preparation leads to:
- Irrelevant retrievals: System finds wrong information
- Context loss: Important details get separated
- Slow performance: Inefficient chunking slows retrieval
- Inaccurate responses: LLM works with incomplete context
- High costs: Wasted tokens and compute resources
Good preparation = Better retrieval = More accurate responses
Step 1: Convert to Clean Markdown
Why Markdown for RAG?
- Semantic structure: Headings help identify topic boundaries
- Token efficiency: Minimal syntax overhead
- Easy parsing: Simple to chunk and process
- Metadata preservation: Structure aids retrieval
- Human readable: Easy to debug and maintain
Conversion Best Practices
- Start with quality sources: Use our ePub converter or PDF converter
- Preserve structure: Maintain heading hierarchy
- Clean artifacts: Remove page numbers, headers, footers
- Fix formatting: Ensure consistent markdown syntax
- Validate output: Check for conversion errors
Step 2: Add Metadata and Frontmatter
Metadata dramatically improves retrieval accuracy:
--- document_id: atomic_habits_ch1 title: The Surprising Power of Atomic Habits source: Atomic Habits by James Clear chapter: 1 category: personal_development topics: - habit formation - behavior change - compound effect keywords: - atomic habits - 1% improvement - identity change - habit loop date_added: 2024-01-15 content_type: book_chapter reading_level: intermediate estimated_tokens: 2500 --- # The Surprising Power of Atomic Habits [Content begins...]
Essential Metadata Fields
- document_id: Unique identifier
- title: Clear, descriptive title
- source: Original source attribution
- category: Primary topic area
- topics: Specific subjects covered
- keywords: Search terms and concepts
- content_type: Book, article, guide, etc.
Optional but Useful
- date_added: When document was added
- last_updated: Most recent modification
- author: Original author
- reading_level: Complexity indicator
- estimated_tokens: Size for context planning
- related_docs: Links to related content
Step 3: Chunking Strategy
Chunking is the most critical step in RAG document preparation. Get this wrong and everything else fails.
What is Chunking?
Breaking large documents into smaller pieces (chunks) that can be:
- Embedded as vectors
- Retrieved independently
- Fit within context windows
- Provide complete, coherent information
Chunking Strategies Compared
1. Fixed-Size Chunking
Method: Split every N tokens/characters
Chunk 1: [0-500 tokens] Chunk 2: [500-1000 tokens] Chunk 3: [1000-1500 tokens]
Pros:
- Simple to implement
- Predictable chunk sizes
- Fast processing
Cons:
- Breaks mid-sentence or mid-concept
- Loses semantic coherence
- Poor retrieval accuracy
Best for: Quick prototypes only
2. Semantic Chunking (Recommended)
Method: Split at natural boundaries (headings, paragraphs, topics)
Chunk 1: ## Introduction + full section Chunk 2: ## Key Concept 1 + full section Chunk 3: ## Key Concept 2 + full section
Pros:
- Preserves semantic meaning
- Natural topic boundaries
- Better retrieval accuracy
- More coherent context
Cons:
- Variable chunk sizes
- Requires document structure
- More complex implementation
Best for: Most RAG applications
3. Sliding Window Chunking
Method: Overlapping chunks with shared context
Chunk 1: [0-500 tokens] Chunk 2: [400-900 tokens] (100 token overlap) Chunk 3: [800-1300 tokens] (100 token overlap)
Pros:
- Reduces context loss at boundaries
- Improves retrieval recall
- Handles concepts spanning boundaries
Cons:
- Increases storage requirements
- More embeddings to compute
- Potential redundancy in results
Best for: Critical applications where accuracy matters most
4. Hierarchical Chunking
Method: Multiple levels of granularity
Level 1: Entire document summary Level 2: Chapter summaries Level 3: Section content Level 4: Paragraph details
Pros:
- Retrieves at appropriate level
- Can provide broad or specific context
- Excellent for large documents
Cons:
- Complex to implement
- Requires careful hierarchy design
- More storage and processing
Best for: Large knowledge bases with complex structure
Recommended Chunk Sizes
Quick facts, definitions, simple answers
Most use cases, balanced context and specificity
Complex topics, detailed explanations
Too large, reduces retrieval precision
Chunking Implementation Example
Semantic chunking by heading level:
# Document Title (not chunked separately) ## Chapter 1: Introduction [Full section content = Chunk 1] ### Subsection 1.1 [If chapter too long, subsection = Chunk 2] ### Subsection 1.2 [Subsection = Chunk 3] ## Chapter 2: Main Content [Full section = Chunk 4]
Step 4: Enhance Chunks with Context
Problem: Context Loss
When you chunk, you lose surrounding context:
Original: # Atomic Habits ## Chapter 1 ### The Four Laws 1. Make it obvious Chunk becomes just: "1. Make it obvious" (Missing: book title, chapter, section)
Solution: Context Enrichment
Add hierarchical context to each chunk:
--- document: Atomic Habits by James Clear chapter: 1 - The Surprising Power of Atomic Habits section: The Four Laws of Behavior Change subsection: Law 1 - Make It Obvious --- # Law 1: Make It Obvious The first law of behavior change is to make it obvious... [Full subsection content] --- Related concepts: habit loop, cue, environment design Previous: Introduction to Four Laws Next: Law 2 - Make It Attractive
Context Enrichment Strategies
- Hierarchical breadcrumbs: Book → Chapter → Section
- Related concepts: Link to related chunks
- Navigation hints: Previous/next sections
- Summary: Brief section summary
- Keywords: Important terms in this chunk
Step 5: Generate Embeddings
What are Embeddings?
Vector representations of text that capture semantic meaning. Similar concepts have similar vectors.
Embedding Model Selection
OpenAI text-embedding-3-large
3072 dimensions, excellent quality
Best for: High-accuracy applications, budget available
OpenAI text-embedding-3-small
1536 dimensions, good balance
Best for: Most applications, cost-conscious
Sentence Transformers (open-source)
384-768 dimensions, free
Best for: Self-hosted, privacy-sensitive
Embedding Best Practices
- Consistent model: Use same model for all documents
- Include metadata: Embed title + content together
- Batch processing: Generate embeddings in batches for efficiency
- Cache embeddings: Store vectors, don't regenerate
- Version control: Track which embedding model version used
Step 6: Optimize for Retrieval
Retrieval Strategies
1. Semantic Search (Vector Similarity)
Find chunks with similar meaning to query:
- Cosine similarity between query and chunk embeddings
- Return top K most similar chunks
- Works well for conceptual queries
2. Keyword Search (BM25)
Traditional keyword matching:
- Exact term matching with relevance scoring
- Works well for specific terms or names
- Fast and efficient
3. Hybrid Search (Recommended)
Combine semantic and keyword search:
Results = (0.7 × Semantic Score) + (0.3 × Keyword Score)
- Best of both worlds
- Handles both conceptual and specific queries
- More robust retrieval
Retrieval Optimization Techniques
- Query expansion: Rephrase query multiple ways
- Reranking: Use LLM to rerank initial results
- Metadata filtering: Pre-filter by category, date, etc.
- MMR (Maximal Marginal Relevance): Diversify results
- Parent-child retrieval: Retrieve small chunks, return larger context
Step 7: Test and Iterate
Create Test Queries
Build a test set of queries and expected results:
Test Cases: 1. Query: "How do I build a new habit?" Expected: Chunks from Atomic Habits, Four Laws section 2. Query: "What's the two-minute rule?" Expected: Specific section on two-minute rule 3. Query: "Compare habit stacking and implementation intentions" Expected: Multiple chunks from different sections
Evaluation Metrics
- Precision@K: How many of top K results are relevant?
- Recall: Did we retrieve all relevant chunks?
- MRR (Mean Reciprocal Rank): Position of first relevant result
- NDCG: Normalized discounted cumulative gain
- Human evaluation: Manual review of response quality
Common Issues and Fixes
Problem: Irrelevant Results
System retrieves wrong chunks
Solutions:
- • Improve chunk boundaries
- • Add more metadata
- • Use hybrid search
- • Implement reranking
Problem: Incomplete Context
Retrieved chunks lack necessary info
Solutions:
- • Increase chunk size
- • Add overlapping chunks
- • Include hierarchical context
- • Retrieve parent chunks
Advanced Techniques
1. Multi-Vector Retrieval
Generate multiple embeddings per chunk:
- One for main content
- One for summary
- One for questions it answers
2. Hypothetical Document Embeddings (HyDE)
Generate hypothetical answer to query, embed that, then search:
- User asks: "How do I build a habit?"
- LLM generates hypothetical answer
- Embed hypothetical answer
- Search for similar real chunks
3. Query Decomposition
Break complex queries into sub-queries:
Complex Query: "Compare habit stacking and implementation intentions, and explain when to use each" Decomposed: 1. "What is habit stacking?" 2. "What are implementation intentions?" 3. "When should I use habit stacking?" 4. "When should I use implementation intentions?"
Tools and Frameworks
Vector Databases
- Pinecone: Managed, easy to use, scales well
- Weaviate: Open-source, feature-rich
- Qdrant: Fast, good for production
- ChromaDB: Simple, good for prototyping
- FAISS: Facebook's library, self-hosted
RAG Frameworks
- LangChain: Most popular, extensive features
- LlamaIndex: Specialized for RAG, excellent docs
- Haystack: Production-ready, modular
- txtai: Lightweight, easy to start
Production Checklist
Before deploying your RAG system:
- ✓ Documents converted to clean markdown
- ✓ Comprehensive metadata added
- ✓ Semantic chunking implemented
- ✓ Context enrichment applied
- ✓ Embeddings generated and stored
- ✓ Hybrid search configured
- ✓ Test queries passing
- ✓ Evaluation metrics tracked
- ✓ Monitoring and logging set up
- ✓ Update process documented
Conclusion
Document preparation is the foundation of successful RAG systems. Invest time in proper conversion, chunking, and enrichment—it directly impacts your system's accuracy and usefulness.
Start with semantic chunking at 400-800 tokens, add comprehensive metadata, and test thoroughly. Iterate based on real-world performance.
Prepare Documents for Your RAG System
Convert your documents to clean, structured markdown optimized for RAG retrieval.