Convert PDF to clean Markdown.
Native PDFs convert instantly in your browser. Scans and complex layouts get an AI boost. Headings, tables and reading order preserved.
Your files are never stored: documents convert in your browser where possible, and anything sent to our AI travels encrypted (HTTPS) and is discarded after processing. Privacy Policy
PDF is the hardest format to convert well: text arrives as positioned fragments with no structure. CleanMD rebuilds the document — it detects heading levels from typography, reconstructs paragraphs and lists, and keeps tables as real Markdown tables. Native PDFs are processed entirely client-side, so your documents never leave your device unless you choose the optional AI pass for complex-layout pages.
Working with related formats? See also Images to Markdown, DOCX to Markdown, PPTX to Markdown.
Why PDF to Markdown is hard
A PDF doesn't know it contains a heading. It stores glyphs at coordinates on a page, and a converter has to guess the structure back from font sizes and positions. Most tools recover the paragraphs and stop there, which is how you end up with page numbers in the middle of sentences and every heading flattened into body text.
CleanMD ranks the font sizes used in the document and maps the larger ones to heading levels. It also looks for bold lines at body size, because that's how LaTeX papers mark their subsections, and when section numbers disagree with the fonts the numbers win, so 2.1 always ends up under 2. On RFC 9110, the 194-page HTTP specification in our public benchmark, all 291 numbered sections come out at the right depth. On the Supreme Court's Loper Bright opinion, 28 of the 31 part markers become headings. Those are the bare I, II, A and B lines that other converters merge into the paragraph below.
Running headers and footers such as "Standards Track [Page 12]" are removed, and so are page numbers. Text set in a monospace font becomes a code block, and a listing that runs across a page break stays one block, which is why the ABNF grammar in RFC 9110 comes out whole. Tables are rebuilt from the position of each word, or from the drawn grid when the table has borders. When the structure isn't clear enough the text is left as text and you get a warning. A wrong table is worse than none.
Scanned PDFs have no text layer, so they go to an OCR engine hosted in the EU with zero data retention, free up to 10 pages. Native PDFs never leave your browser. Some things still go wrong and we'd rather say so: math inside a paragraph loses its subscripts, and two tables printed side by side in a two-column paper can merge into one.
How it works
- 1Drop your PDF in the converter above, or click to browse.
- 2The converter detects whether the PDF is native text or scanned, and picks the fastest engine automatically.
- 3Preview the Markdown with correct H1/H2/H3 hierarchy, then download the .md file.
Before and after
QUARTERLY SAFETY REPORT The Occupational Safety and Health Act of 1970 was passed to prevent workers from being killed or seriously harmed at work. Key Requirements • Employers must provide a safe workplace • The Act created OSHA Department Incidents Risk Logistics 2 Moderate Warehouse 0 Low
# Quarterly Safety Report The Occupational Safety and Health Act of 1970 was passed to prevent workers from being killed or seriously harmed at work. ## Key Requirements - Employers must provide a safe workplace - The Act created OSHA | Department | Incidents | Risk | | ---------- | --------- | -------- | | Logistics | 2 | Moderate | | Warehouse | 0 | Low |
Frequently asked questions
Is my PDF uploaded to a server?
Native PDFs are converted entirely in your browser — the file never leaves your device. The one exception is opt-in: if some pages have a complex layout (side headings, narrow columns), the result offers to re-process just those pages with the OCR partner below, and nothing is sent unless you accept. Scanned PDFs need real OCR, so they are processed server-side, primarily by our EU-based OCR partner (Mistral AI) under a zero-data-retention agreement: your document is never stored, never logged beyond producing the result, and never used to train models. On every tier, nothing is kept after the conversion.
Does it work with scanned PDFs?
Yes. The converter detects scanned pages automatically and routes them through a layout-aware OCR engine, so you still get structured Markdown — headings, paragraphs and tables — instead of a wall of unformatted text.
Are headings and tables preserved?
Yes — that's our specialty. Heading levels (H1/H2/H3) are inferred from the document's typography and section order, and tables are emitted as real Markdown tables, not tab-separated text.
Is PDF to Markdown free?
Yes for most PDFs: 3 conversions a day with no account, and PDFs can be up to 50 MB. Scanned PDFs are free up to 10 pages. A longer scan shows a 3-page preview and then unlocks with a one-off payment of $3 (up to 20 pages) or $4 (up to 100), no subscription needed.