What makes a scan different
A scanned PDF is a stack of photographs. There's no text in it to select, so the usual PDF tools give you an empty note or, worse, the garbage text layer some scanners add on their own. Getting a usable Obsidian note out of it takes OCR first, and then the harder part, which is putting the structure back: which lines are headings, where a paragraph continues on the next page, which rows belong to the same table.
CleanMD checks every page in your browser before anything is sent. A page with no text and a full-page image counts as scanned. If no page in the file has real text, the whole PDF goes to OCR. That includes PDFs where the scanner already embedded its own OCR layer, because that layer is often bad enough to be worse than starting over. We found one with 6,000 characters per page of mangled text, which is why the check looks at the pages themselves rather than trusting the character count.
What the OCR keeps and what it drops
Scans up to 10 pages are free. Under 25 MB they're read by Mistral's OCR, processed in the EU with zero data retention. A larger file, or any scan while that service is unavailable, is read by Google's Gemini instead. After the OCR the text goes through the same cleanup as everything else. Running headers and footers at the top and bottom of each page are removed. A paragraph that breaks across two pages is joined back together, and a word hyphenated across the break is mended. A table that continues on the next page becomes one table, without the repeated header row. Heading levels are tidied so they never skip, so a ### never sits directly under a #.
Figures are handled with less certainty. On Premium, with the figure option left on, each one is cropped and sent to Google's Gemini for a description. When it can be described, the description goes in a quoted block that starts with "Image interpretation" and the page number, so you always know it was written by a model and not copied from the page. On the free plan, or with the option off, or when a figure can't be described, it's left out. Equations are the weakest part of a scan: there's no special handling for them on this path, so check any formula before you rely on it.
A scan longer than 10 pages shows a 3-page preview first. The full document then costs a one-off $3 for up to 20 pages or $4 for up to 100, with no subscription, and scans over 100 pages aren't supported. Longer scans, preview included, are read as page images by Google's Gemini instead of Mistral, and the provider deletes the files within 48 hours.
The one case that catches people out
A PDF that mixes typed pages and scanned pages is treated as a normal PDF, because sending it to OCR would also send the pages that could have stayed in your browser. The typed pages convert locally and the scanned ones are skipped, with a warning that says how many. If those pages matter, split them into their own PDF and convert that file separately. It will then go through OCR like any other scan.
Getting it into your vault
The download is a .md file named after the PDF, so contract-2019.pdf becomes contract-2019.md, and you can drop it straight into the vault folder. It opens with a YAML block that Obsidian shows as properties: the source file name and the format, plus the title when one is found. Copy gives you exactly the same text, properties included.
Because the headings are real Markdown headings, the outline pane works from the start and you can link to a section with [[contract-2019#Payment terms]] like any note you wrote yourself. Tables are standard pipe tables, which Obsidian renders without a plugin. Images are not saved as attachments. What survives of a figure is its description, which keeps the vault light but means a diagram you need to see should be clipped from the original.
If you're scanning the document yourself, scan it to one PDF rather than a set of photos. The OCR then sees the pages in order and keeps the heading levels consistent from page to page, which separate images can't give it.
Converters for this
Last updated 2026-10-05. More guides: Markdown for RAG · Word with tracked changes · YouTube lecture to notes