Convert PDF to Markdown
Convert a PDF into clean Markdown with real headings, lists and tables — the format for feeding a document to an AI model or a static site.
Drop your file here
Processed securely and deleted within the hour.
Why this format, and why now
Markdown has become the way documents get fed to things. An AI model reads it far better than a PDF, because the headings and lists are explicit rather than implied by font size. A static site generator takes it directly. A wiki, a docs repository, a note-taking app, a diff — all of them want text with structure, and none of them want a PDF.
The hard part is not producing text. PDF to text already does that, and gives you a wall of paragraphs. The hard part is recovering the structure, because a PDF does not know it has headings. It has characters, at coordinates, in fonts, at sizes. Everything else is inferred.
How the headings are worked out
The body text size is the modal font size across the document — the size most of the characters are actually set in. Not the average, which a title drags up and footnotes drag down.
Every size above that is then collected, sorted, and ranked: largest becomes h1, next becomes h2, and so on down. Ranking rather than thresholding is what makes this work on real documents. A report set in 10pt with 12pt headings and a report set in 12pt with 24pt headings both need an h2, and no fixed cutoff serves both.
Some documents mark headings with bold at body size instead of a larger size. There is an option for that, off by default, because switched on unnecessarily it promotes every emphasised phrase in the document to a heading.
Tables, and the mistake nearly every converter makes
Detected tables become proper Markdown tables. The part that matters is what happens to the text inside them: the table region is excluded from the flowing text.
Skip that step and every table appears twice — once as a table, and again as a run of paragraphs made of its cells, in reading order, meaning nothing. It roughly doubles the length of the output and it is the single most common defect in this category. A pipe character inside a cell is escaped too, since an unescaped one silently ends the cell and shifts the whole row.
The unmapped glyph problem
A bullet is often set in a symbol font with no character map attached. Extract the text and you get (cid:127) — a glyph index, not a character. This is extremely common, and it is why converted documents so often contain literal (cid:127) where the bullets should be.
A leading unmapped glyph is, overwhelmingly, a bullet. So it is read as one and the line becomes a list item. Unmapped glyphs anywhere else are dropped rather than printed, and the result tells you how many went, because that count is your signal that the PDF has font problems worth knowing about.
Two columns
A two-column paper read straight down the page interleaves the columns line by line. The output looks like words and is not sentences, and it is useless for either reading or feeding to a model.
Real columns have a gutter — a vertical band that no word crosses — and a single-column page does not, so this is a property of the layout rather than a guess. When one is found, the left column is read fully and then the right. You can force either behaviour if the detection gets it wrong on an unusual page.
Code keeps its indentation
Monospaced lines become indented code blocks, and the indentation is recovered from the horizontal position of the first character rather than from leading spaces — because extraction gives you words and their coordinates, not whitespace.
That distinction matters more than it sounds. A naive converter turns
def f():
return 1
into two lines at the same level, which for code is not an untidy result, it is a different program.
Markdown characters in your text are escaped
A document that mentions margin_growth or *rose* or [see note] will, in unescaped Markdown, render as italics, emphasis and a broken link. Those get escaped — and only those. Over-escaping is its own failure: a file full of backslashed punctuation is harder to read than the problem it fixes, so a # in the middle of a sentence is left alone, because it only means anything at the start of a line.
Hyperlinks are kept as real Markdown links. The order matters there too: the markup is attached after escaping, since doing it the other way round escapes the brackets you just wrote and the link renders as literal text.
It will not work on a scan
A scanned page is a photograph. There is no text in it to find and no structure to infer, so you would get an empty file. That is detected and refused with an explanation rather than handed to you as a successful empty result — run OCR over it first to add a text layer, then come back.
Common questions
Will it work on a scanned PDF?
No, and it says so rather than giving you an empty file. A scan is a photograph with no text in it. Run OCR over it first to add a text layer.
How does it decide what is a heading?
The body size is the most common font size in the document. Larger sizes are then ranked — biggest becomes h1, next h2 — so it works whether the document is set in 10pt or 12pt.
Why do other converters leave "(cid:127)" in the text?
Because a bullet in a symbol font has no character map, so it extracts as a glyph index rather than a character. A leading one is read as a bullet here, and the rest are dropped and counted.
Do tables come out as Markdown tables?
Yes, and their cells are excluded from the surrounding text. Without that step every table appears twice — once as a table and again as meaningless paragraphs.
My two-column paper came out scrambled.
Set the column layout to two. Detection looks for a gutter no word crosses, which most papers have, but an unusual layout can defeat it.
Does code keep its indentation?
Yes. It is recovered from the horizontal position of the first character, since extraction gives coordinates rather than leading spaces.
Why are there backslashes in my text?
Characters that would otherwise render as Markdown — underscores, asterisks, square brackets — are escaped so your text reads as written. Only those, and only where they would matter.