This article was generated with AI assistance. While we strive for accuracy, please verify any technical specifications or recommendations independently before implementing them in your AI training pipeline.
Choosing the right format for LLM training data significantly impacts model performance, training efficiency, and downstream application quality. This guide compares all major formats.
Quick Comparison Table
| Format | Best For | Token Efficiency | Structure | Human Readable |
|---|---|---|---|---|
| Markdown | General text, docs, books | Excellent | High | Very High |
| JSON | Structured data, APIs | Moderate | Very High | Moderate |
| Plain Text | Simple content | Excellent | Low | High |
| HTML | Web content | Poor | High | Low |
| XML | Complex hierarchies | Poor | Very High | Low |
| CSV | Tabular data | Excellent | Moderate | High |
Markdown: The Sweet Spot
Why Markdown Excels
Markdown has become the de facto standard for LLM training for good reasons:
1. Token Efficiency
Markdown uses minimal syntax overhead:
# Heading (2 tokens) vs <h1>Heading</h1> (5 tokens in HTML) **bold** (3 tokens) vs <strong>bold</strong> (5 tokens in HTML)
Impact: 30-50% fewer tokens compared to HTML/XML, meaning lower training costs and faster processing.
2. Native LLM Understanding
- Most LLMs (GPT-4, Claude, Llama) are trained extensively on markdown
- Models "think" in markdown naturally
- Better comprehension of structure and hierarchy
- More accurate responses when querying markdown content
3. Human Readability
Critical for:
- Quality control and review
- Debugging training data
- Manual curation and editing
- Collaboration across teams
4. Semantic Structure
Markdown preserves meaning without visual clutter:
- Clear heading hierarchy (H1-H6)
- Lists (ordered and unordered)
- Emphasis (bold, italic)
- Code blocks with syntax highlighting
- Links and references
- Tables and blockquotes
Markdown Limitations
- Complex layouts: Not ideal for multi-column or intricate designs
- Strict schemas: Less suitable than JSON for rigid data structures
- Mathematical notation: Requires extensions (LaTeX)
- Metadata: Needs frontmatter conventions
Best Use Cases for Markdown
- Books and long-form content
- Documentation and guides
- Blog posts and articles
- Educational materials
- Conversational training data
- Knowledge bases
JSON: Structured Data Champion
When JSON Shines
1. Strict Structure Requirements
{
"instruction": "Explain quantum entanglement",
"input": "",
"output": "Quantum entanglement is...",
"metadata": {
"category": "physics",
"difficulty": "advanced",
"tokens": 150
}
}Perfect for:
- Instruction-tuning datasets
- Question-answer pairs
- Structured prompts and completions
- API training data
2. Machine Processing
- Easy to parse programmatically
- Validate with JSON Schema
- Transform and filter efficiently
- Integrate with data pipelines
3. Metadata Rich
- Embed extensive metadata
- Track provenance and quality
- Enable sophisticated filtering
- Support multi-field indexing
JSON Drawbacks for LLM Training
- Token overhead: Brackets, quotes, commas add 20-40% overhead
- Less natural: LLMs don't "think" in JSON
- Hard to read: Difficult for humans to review large files
- Verbose: Repetitive structure wastes tokens
Hybrid Approach: JSON + Markdown
Best of both worlds:
{
"id": "habit_001",
"category": "coaching",
"content": "# Building Better Habits\n\nHabits form through...\n\n## Key Principles\n\n1. **Start small**...",
"metadata": {
"source": "atomic_habits",
"chapter": 1,
"tokens": 450
}
}Use JSON for structure, markdown for content.
Plain Text: Simplicity at a Cost
Advantages
- Maximum token efficiency
- No parsing overhead
- Universal compatibility
- Smallest file sizes
Major Limitations
- No structure: LLM can't distinguish headings from body text
- No emphasis: Important concepts not highlighted
- No metadata: Can't embed context or provenance
- Poor navigation: Difficult for LLM to locate specific sections
When to Use Plain Text
- Simple, unstructured content
- When structure is irrelevant
- Maximum token efficiency required
- Preprocessing into other formats
HTML: Web Content Challenges
Why HTML is Problematic
1. Massive Token Overhead
<div class="content-wrapper">
<article class="post-content">
<h1 class="post-title">My Title</h1>
<p class="post-body">Content here</p>
</article>
</div>
vs
# My Title
Content hereHTML can use 3-5x more tokens for the same content.
2. Visual vs Semantic
- HTML mixes presentation with content
- CSS classes add noise
- Inline styles waste tokens
- JavaScript and scripts are irrelevant
3. Inconsistent Structure
- Every website uses different HTML patterns
- Difficult for LLM to learn consistent structure
- Lots of boilerplate and navigation elements
When HTML Makes Sense
- Training on web scraping tasks
- Teaching HTML generation
- Web-specific applications
- When converting to markdown is impractical
Best Practice: Convert HTML to Markdown
Use our HTML to Markdown converter to clean up web content before training.
XML: Enterprise Data Format
XML Characteristics
Advantages
- Extremely structured and hierarchical
- Schema validation (XSD)
- Namespace support for complex systems
- Industry standard for many domains
Disadvantages for LLM Training
- Verbose: Even worse than HTML for token efficiency
- Redundant: Opening and closing tags repeat information
- Unnatural: LLMs don't naturally work with XML
- Hard to read: Difficult for human review
When to Use XML
- Domain-specific applications (legal, medical)
- When source data is already XML
- Teaching XML parsing or generation
- Strict validation requirements
CSV: Tabular Data
Perfect for Structured Tables
name,age,occupation,location Alice,32,Coach,New York Bob,45,Consultant,London Carol,28,Therapist,Sydney
Strengths
- Extremely token-efficient for tabular data
- Easy to parse and process
- Universal support
- Good for numerical data
Limitations
- Only works for flat, tabular data
- No hierarchy or nesting
- Limited text content per cell
- No formatting or emphasis
Best Use Cases
- Training on data analysis tasks
- Numerical reasoning
- Entity databases
- Simple fact tables
Specialized Formats
LaTeX
Best for: Mathematical and scientific content
- Precise mathematical notation
- Academic papers
- Technical documentation
Drawback: Verbose and complex syntax
YAML
Best for: Configuration and metadata
- More human-readable than JSON
- Good for frontmatter
- Hierarchical data
Drawback: Whitespace-sensitive, less common in LLM training
Parquet/Arrow
Best for: Large-scale data processing
- Columnar storage
- Efficient compression
- Fast querying
Drawback: Binary format, not human-readable
Real-World Performance Comparison
Token Count Example
Same content in different formats:
Clean, structured, efficient
+22% overhead from structure
+89% overhead from tags and classes
+104% overhead from verbose syntax
Training Cost Impact
For 1 million training examples:
- Markdown: 450M tokens = $3,600 (baseline)
- JSON: 550M tokens = $4,400 (+22%)
- HTML: 850M tokens = $6,800 (+89%)
- XML: 920M tokens = $7,360 (+104%)
*Estimated costs based on typical LLM training pricing
Recommendations by Use Case
Use Markdown For:
- ✓ Books and long-form content
- ✓ Documentation and guides
- ✓ Blog posts and articles
- ✓ Conversational AI training
- ✓ Knowledge bases
- ✓ Educational materials
- ✓ General-purpose LLM training
Use JSON For:
- ✓ Instruction-tuning datasets
- ✓ Q&A pairs
- ✓ API training data
- ✓ Structured prompts/completions
- ✓ Metadata-heavy datasets
- ✓ Programmatic processing
- ✓ Fine-tuning pipelines
Hybrid Strategies
Strategy 1: JSON Wrapper + Markdown Content
Combine structure with efficiency:
{
"id": "doc_001",
"type": "book_chapter",
"source": "atomic_habits",
"chapter": 1,
"content_format": "markdown",
"content": "# The Surprising Power of Atomic Habits\n\nThe backbone of this book is my four-step model of habits..."
}Strategy 2: Markdown + YAML Frontmatter
Metadata in YAML, content in markdown:
--- title: Atomic Habits - Chapter 1 author: James Clear type: book_chapter category: personal_development tags: [habits, behavior-change] --- # The Surprising Power of Atomic Habits The backbone of this book is my four-step model...
Strategy 3: Format-Specific Datasets
- Markdown: 70% of training data (general content)
- JSON: 20% (structured tasks)
- CSV: 5% (tabular data)
- Code: 5% (programming tasks)
Conversion Best Practices
From HTML to Markdown
- Strip all styling and scripts
- Convert semantic tags (h1-h6, p, ul, ol)
- Preserve links and images
- Remove navigation and boilerplate
- Clean up whitespace
From PDF to Markdown
- Extract text with structure preservation
- Identify heading hierarchy
- Maintain lists and formatting
- Handle tables appropriately
- Extract and reference images
Use our PDF to Markdown converter for optimal results.
From ePub to Markdown
- Parse ePub structure (chapters, sections)
- Convert XHTML to markdown
- Preserve semantic meaning
- Handle embedded media
- Maintain reading order
Conclusion
For most LLM training use cases, markdown is the optimal format. It balances token efficiency, human readability, structural clarity, and native LLM understanding.
Use JSON when you need strict structure or extensive metadata. Use plain text only for simple, unstructured content. Avoid HTML and XML unless absolutely necessary.
The best approach is often hybrid: JSON for structure and metadata, markdown for content. This gives you the benefits of both formats while minimizing their drawbacks.
Convert Your Content to Markdown
Get token-efficient, AI-ready markdown from any document format.