Convert HTML to clean Markdown.
h1–h6, links, lists, tables and code blocks mapped to proper Markdown. No boilerplate, no inline styles.
Your files are never stored: documents convert in your browser where possible, and anything sent to our AI travels encrypted (HTTPS) and is discarded after processing. Privacy Policy
HTML is the closest cousin to Markdown, but real-world pages are full of wrappers, navigation and inline styles. CleanMD keeps the semantic content — headings, paragraphs, links, lists, tables, code blocks — and drops the noise, producing Markdown that reads like it was written by hand.
Working with related formats? See also URL to Markdown, DOCX to Markdown, CSV to Markdown.
Keeping the article, dropping the page around it
An HTML file carries a lot besides its content. Navigation bars, cookie banners, scripts, styles and footers full of links all sit next to the text you actually want, and a converter that keeps everything gives you a Markdown file where the article starts halfway down.
CleanMD removes scripts, styles and navigation, and when the page marks its main content with a main or article element it converts only that. Headings keep their levels, code blocks keep their language where the page declares it, tables become Markdown tables and links keep their targets. Content hidden inside a collapsed details element is kept, because hidden on screen doesn't mean unimportant.
Headings get cleaned on the way out as well. A line break inside a heading would split the Markdown title in two, so it becomes a space. Bold wrapped around a whole heading is dropped, since a heading is emphasis already, and icons inside headings are removed so they don't end up in the outline.
Forms are a harder call than they look. On a documentation page or a style guide the form is the content, so fieldsets keep their legend, buttons keep their label, the options of a select are listed with the chosen one in bold, and checkboxes become [ ] and [x]. Search boxes and newsletter sign-ups are still thrown away, since those are part of the site rather than the page. Password fields and hidden inputs are never copied.
One limit we know about: when two pieces of text are spaced only by CSS, for example items in a flex row, they come out joined without a space. The converter sees the HTML, not the layout, and there's no way to tell those apart from text that really is continuous. The file is converted in your browser and isn't uploaded.
How it works
- 1Drop an .html file in the converter above.
- 2Semantic tags (h1–h6, ul/ol, table, pre/code, a) are mapped to their Markdown equivalents; boilerplate is dropped.
- 3Preview the result and download the .md file.
Before and after
<article>
<h1>Getting started</h1>
<p>Install the CLI with <code>npm i -g tool</code>.</p>
<h2>First steps</h2>
<ul>
<li>Create a <a href="/docs/project">project</a></li>
<li>Run <code>tool dev</code></li>
</ul>
</article># Getting started Install the CLI with `npm i -g tool`. ## First steps - Create a [project](/docs/project) - Run `tool dev`
Frequently asked questions
Does it keep links and code blocks?
Yes. Anchors become [text](url) links, <pre>/<code> become fenced code blocks and inline code, and tables become Markdown tables.
What happens to navigation menus and scripts?
Non-content elements (scripts, styles, boilerplate) are dropped so the Markdown contains only the readable content.
Can I convert a live URL?
Yes — paste the page URL in the converter (or use the URL to Markdown page) and it converts for free, no file needed.
Is the file uploaded?
No. HTML conversion runs entirely in your browser and nothing is stored.