Convert PDF to text
Pull the text out of a PDF.
Drop your file here
It is processed on your device — nothing is uploaded.
Extraction, not recognition
This reads the text layer a PDF already contains. Every PDF made from a word processor, a web page or a report generator has one: the characters are stored as characters, and pulling them out is exact.
What it cannot do is read a photograph. If your PDF came from a scanner or a phone camera, the pages are images and there is no text to extract — only pixels arranged to look like letters. The tool detects this and tells you plainly rather than handing back an empty file, which is the failure mode that leaves people assuming the tool is broken.
Layout mode
By default you get the text in reading order, cleanly. Turning on "keep the page layout" preserves the horizontal positioning with spaces, which reconstructs columns and simple tables as they appeared on the page. It is the right choice when position carries meaning and the wrong one when you just want prose, because it introduces a lot of whitespace.
Where reading order goes wrong
A PDF stores marks at coordinates, not paragraphs, so reading order has to be inferred. Single-column documents are reliable. Multi-column layouts sometimes interleave the columns, and text in sidebars or captions can arrive in unexpected places. This is inherent to the format rather than specific to this tool.
What gets lost
Bold, italics, headings, colours and images. Plain text keeps none of it. If you want the formatting, PDF to Word preserves paragraphs, headings and tables.
Why plain text is still useful
It is the format everything else accepts — search, scripts, spreadsheets, language models. For pulling data out of documents at any scale, plain text is usually the right destination.
Common questions
Why is my output empty?
The PDF is a scan, so it contains images rather than text. Extraction cannot read pixels — that needs OCR.
Does it keep formatting?
No, this is plain text. Use PDF to Word if you need headings, bold and tables.
What does layout mode do?
It preserves horizontal positioning with spaces, reconstructing columns and simple tables. Useful when position matters, noisy when it does not.
Why is the text out of order?
A PDF stores marks at coordinates rather than paragraphs, so reading order is inferred. Multi-column layouts are the usual culprit.
Can I extract only some pages?
Yes, the pages field accepts ranges.
Is it accurate?
Exact, where a text layer exists — the characters are copied, not recognised.