Why the conversion step matters more than it looks
A common way to chunk documents for retrieval is to split on headings, with something like LangChain's MarkdownHeaderTextSplitter, and keep the heading path as metadata on every chunk. A chunk that knows it sits under "Conditional Requests > If-Match" is far easier to retrieve and to cite than the same paragraph floating on its own. That only works if the converter produced real # and ## lines in the first place.
PDFs are where this breaks. A PDF has no headings, only text in larger or bolder fonts, and many converters hand you every line as a paragraph. The splitter then has nothing to split on, and you fall back to fixed-size chunks that cut sections in half. The second problem is noise. A running footer such as "Standards Track [Page 12]" repeated on every page ends up inside your chunks, gets embedded, and can match queries it has nothing to do with.
What we measured
We publish a benchmark on five public PDFs, scored against independent ground truth such as each document's own table of contents or an official text edition, with the script in a public repository so anyone can rerun it. The run of 15 September 2026 compared CleanMD, MarkItDown and pymupdf4llm on default settings.
On RFC 9110, the 194-page HTTP specification, CleanMD found all 291 numbered sections as headings, all at the right depth, and left no running footer lines in the text. pymupdf4llm found 290 sections, 275 of them at the right depth, and kept 194 footer lines. MarkItDown produced no headings at all and 187 footer lines. Its own README says its output is meant for text analysis tools, so this is a design choice and not a bug, but for header-based chunking it leaves you nothing to split on.
Code matters too if your corpus is technical. Text in a monospace font becomes a fenced code block, so a listing stays in one piece instead of being reflowed into a paragraph. In the same RFC the collected ABNF grammar comes out as a single block, and Think Python, a 292-page book, comes out with 569 fenced blocks.
Where CleanMD is not the right tool
Tables are our weak spot. On the NIST Cybersecurity Framework, 23 table cells in our output contained a whole sentence of prose, against 12 for pymupdf4llm and none for MarkItDown. Some of those are long cells the original table really has, but not all of them. When the structure of a table isn't clear we leave the text as text and say so, but a dense layout can still fool it.
Volume is the other limit. CleanMD converts one document at a time on a web page, with no batch API. If you're ingesting thousands of files, a Python library such as MarkItDown or pymupdf4llm belongs in the pipeline, and the benchmark shows what you trade for that. CleanMD fits the documents where structure decides whether retrieval works: the long specification, the contract, the manual someone will ask detailed questions about.
Small things that help downstream
Every file starts with a YAML block holding the title, the source file name and the format, which you can copy into chunk metadata so an answer can point back to the document it came from. Heading levels never skip, so the heading path on a chunk is always a real path. Repeated page numbers, headers and footers are removed before you see the file, which is one less cleaning step in your own code.
Scanned PDFs go through OCR with the same cleanup, so a scanned manual gives you headings to split on as well. Word files keep their heading styles. Web pages lose their navigation menus and, when the page marks where its main content is, everything around the article. All of these give the splitter the same kind of input.
Converters for this
Last updated 2026-10-05. More guides: Scanned PDF to Obsidian · Word with tracked changes · YouTube lecture to notes