Markdown for RAG: Why Format Decides Retrieval Quality

Everyone tunes chunking, embeddings, and reranking. Almost nobody tunes the one step upstream that caps all three.

Markdown Converters team
Updated Jul 2, 2026
13 min read
Documents flowing through conversion into clean, structured Markdown chunks that feed a vector database knowledge base.

Your RAG assistant answers a question about the Q2 numbers. It's specific, it's confident, and the figure it quotes belongs to a different region. You dig into the retrieval trace: the right chunk came back. The problem is what that chunk contained, a column of bare numbers with no rows attached, because the source PDF's table got flattened into text the moment it was ingested, months ago, before a single embedding was computed.

You'll spend the next week tuning the retriever. Different chunk sizes, a reranker, a better embedding model. None of it will fix this, because the information the model needed was destroyed upstream, at conversion, and every layer after that is just retrieving the wreckage more precisely.

Document format is the most under-examined variable in RAG. It decides whether the chunk your retriever nails is actually usable, and it does so silently, before any of the parts people obsess over even run. This guide is about that step: why Markdown is the format that keeps RAG honest, and what "good conversion" actually means, backed by measurements rather than the unsourced percentages that float around this topic.

We'll go from the failure you can't see, to why Markdown specifically, to how to choose a converter, to getting it right without over-engineering it.

The layer of your RAG stack nobody tunes

Open any RAG optimization guide and you'll find the same list: chunk size, chunk overlap, embedding model, top-k, reranking, query rewriting. All real levers. All downstream of a step the guides skip in a sentence: "first, extract the text."

That sentence is doing enormous work. "Extract the text" is where a table becomes either a clean grid or a pile of orphaned numbers, where a document's section structure is either preserved or erased, where a scanned page is either read or silently dropped. Everything the fancy levers do, they do to whatever came out of that step. Garbage in isn't a cliché here, it's the architecture.

I think this gets skipped because it feels like plumbing, not machine learning. Tuning an embedding model is interesting. Choosing a document converter is boring. But the boring step sets the ceiling on everything the interesting steps can achieve, and a low ceiling can't be raised by polishing the floor.

Architecture rule: retrieval quality has a hard ceiling set by conversion quality. Tune conversion first, because nothing downstream can recover what it discards.

What actually happens to a document, and where format bites

A document takes four steps to become something a RAG system can use, and format matters at three of them.

  1. 1. Convert: the file becomes text. This is where structure is kept or lost. It's the step this article is about.
  2. 2. Chunk: the text is split into pieces small enough to embed. Good chunks follow the document's real boundaries; bad chunks cut mid-table or mid-sentence.
  3. 3. Embed: each chunk becomes a vector. Cleaner, more coherent text produces embeddings that match the right queries.
  4. 4. Retrieve: at query time, the closest chunks are pulled into the model's context.

Notice that steps 2 and 3 inherit whatever step 1 produced. If conversion kept the heading structure, chunking can split on sections and embedding sees coherent topics. If conversion flattened everything into a wall of text, chunking falls back to cutting every N characters, and a table can get sliced down the middle into two meaningless halves in two different vectors.

The step everyone starts optimizing at, retrieval, is step four. Three chances to preserve or destroy the information have already passed.

The failure you can't see: a chunk that retrieved but can't be read

Here's what makes format failures so hard to catch. Retrieval works. The right chunk comes back with a high similarity score. Your evaluation harness, if it measures retrieval by whether the correct chunk was returned, says everything is fine. And the answer is still wrong, because the returned chunk is unreadable.

The same query hits the vector store and returns a chunk. From a flattened PDF the chunk is a column of orphaned values (Region, Q1 revenue, APAC, 1988410) and the model can't tell which number belongs to which row, producing a wrong answer. From clean Markdown the chunk is an intact pipe table, so the model reads it correctly.

This isn't hypothetical. We ran a realistic report through several converters and measured what survived. Plain PDF text extraction, the default for most "load the document" steps, kept zero heading markers and turned both data tables into loose columns of numbers. A layout-aware converter kept every heading and both tables as clean pipe tables. Same document, same information going in; one output the model can reason over, one it can only guess against. The full method and numbers are in our converter benchmark and the three-way comparison.

So the chunk in the bad path retrieves perfectly and answers wrongly, and it does it consistently, on every query that touches that table, forever, until someone re-ingests the source. That's the quiet tax of skipping conversion quality: not a loud error you catch in testing, but a steady wrongness you only notice when a user does.

Why Markdown specifically

"Just keep the structure" is the goal; Markdown is the format that hits it without the costs of the alternatives. Three properties make it the right target for a RAG pipeline.

Headings give chunkers real boundaries

A Markdown document carries an explicit outline: #, ##, ###. That's not decoration, it's a machine-readable table of contents, and the RAG ecosystem already treats it as one. LangChain ships a Markdown header text splitter that cuts chunks on heading boundaries and attaches the section titles as metadata; LlamaIndex has the equivalent. Feed those splitters plain flattened text and they have nothing to split on, so they fall back to counting characters. Feed them Markdown and every chunk lands on a topic boundary and knows which section it came from.

Tables survive as tables

A Markdown pipe table keeps the relationship between a value and its row and column in plain text a model reads natively. That's the exact thing raw extraction destroys, and it's the difference between the two chunks in the diagram above. For any document where the answers live in tables, financial reports, spec sheets, comparison matrices, this property alone justifies the conversion step.

It stays token-lean, which matters per chunk

Markdown carries structure at a fraction of HTML's token cost, about 25% fewer tokens than the raw HTML of the same document in our token measurement. In RAG that compounds: leaner chunks mean you can fit more relevant context in the same window, or pack more meaning into the same chunk-size budget. One honest caveat from that same measurement, against plain PDF text, Markdown costs slightly more tokens, because it's keeping the tables. That's the right trade for retrieval: a few tokens for a chunk that's actually correct.

Working advice: don't chase raw token minimization in a RAG pipeline. Chase structured chunks. A slightly larger chunk that preserves a table beats a smaller one that scrambled it.

Choosing a converter for RAG

The right converter depends on your documents, not on GitHub stars. Match the tool to your hardest input.

Your documentsReach forWhy
Clean, born-digital PDFs & Office filesA fast layout-aware parser, or a hosted converterStructure is easy here; don't pay for heavy models
Complex layouts, dense tables, papersA heavyweight tool (Docling, Marker)Only these recover hard structure reliably
Scans, photos, image PDFsVision-based OCRNo text layer exists; it must be recovered from pixels
A mix, and no time to wire librariesA hosted converter with OCROne path handles clean, complex, and scanned

The one converter to avoid for a RAG pipeline built on PDFs is anything using plain text extraction, which is what popular general tools like MarkItDown do for PDFs. It's excellent for Office files and wrong for the tables your retriever will be asked about. Our alternatives guide maps this by use case.

If you'd rather not maintain a parsing stack, our converter and API target clean, chunk-ready Markdown and run vision OCR on scans, so the same ingestion path covers the whole document mix.

Getting it right in practice

Once conversion is solid, the rest of the pipeline gets easier, and a few habits keep it that way.

  • Convert first, then chunk on headings. Use a Markdown-aware splitter so chunks align with sections and carry their heading as metadata. It improves both retrieval and the model's sense of where an answer came from.
  • Keep tables whole. Set your chunker to avoid splitting inside a table; a half-table in a chunk is worse than the table in its own slightly-oversized chunk.
  • Spot-check the converted Markdown, not just the answers. Read a few converted files before they're embedded. Bad conversion is obvious to a human in seconds and invisible in an eval that only checks retrieval hit-rate.
  • Re-ingest when you improve conversion. Better retrieval settings apply instantly; better conversion only helps documents ingested after the change. Plan a re-index.

The step-by-step version of the full pipeline, chunking strategy, metadata, embeddings, and testing, is in our 7-step RAG preparation workflow. This article is the argument for spending your first hour on the conversion step; that one is the how-to for everything after it.

Keep the first version narrow. Pick one document type, get its conversion clean, verify the chunks by eye, and only then widen. A RAG system that's correct on a small, well-converted corpus is worth more than one that's plausible on a large, flattened one.

Frequently asked questions

What's the best document format for RAG?

Clean Markdown. It preserves headings and tables, gives chunkers natural boundaries, and stays token-lean. Raw text loses structure, HTML adds overhead, and images are expensive and unreliable.

Does format really affect RAG accuracy?

It caps it. A table flattened at conversion is stored broken in your vector database, and no retriever, embedder, or reranker can rebuild what conversion discarded. We measured default extraction dropping every table and heading on a normal report.

How does Markdown improve chunking?

Its heading hierarchy lets structure-aware splitters cut chunks on section boundaries and tag them with their section, so chunks match topics instead of arbitrary character counts. LangChain and LlamaIndex both ship Markdown header splitters for exactly this.

Should I convert scanned PDFs for RAG?

Always. A scan has no text layer, so it embeds as nothing useful. Run vision OCR to produce Markdown first. Skipping this is the top reason a RAG system silently ignores whole documents.

If you're building a RAG pipeline and want the ingestion step to stop being the ceiling, MarkdownConverters turns PDFs, Office files, and scans into clean, chunk-ready Markdown, tables and headings intact, via the browser or the API.

Related reading

Conversion figures measured 2026-07-02 on fixture documents; your corpus will vary. RAG framework APIs change, verify splitter behavior against current docs.