PDF to text

Pull the text out of a PDF in reading order, ready to copy or save as a .txt file. Tells you plainly when a PDF is a scan with no text in it.

Drop a PDF. The text layer is read out in reading order.

Your file is processed on this device. Nothing is uploaded.

When copying from the PDF isn’t enough

Selecting text in a PDF reader works for a paragraph. For a whole document it gets painful: page headers and footers come along, selection jumps between columns and some readers refuse to select across pages at all. Extracting the whole text layer at once gives you everything in one place to search, quote, feed into another tool or paste into a document.

Text PDFs and scanned PDFs

A PDF exported from a word processor, a web page or accounting software contains real text and extraction is exact. A PDF from a scanner or a phone scanning app usually contains images of pages and there is no text inside at all, even though it looks identical on screen. The quick test: if you can’t select a word in your PDF reader, there is nothing to extract.

Using the result

Page markers make it easy to cite a page or check a passage against the original; switch them off for continuous text. The download is plain UTF-8 text, which opens in anything and pastes cleanly into spreadsheets, notes and code. If you need a page as a picture instead,PDF to PNG keeps the layout exactly.

Common questions

Why does my PDF produce no text?

Because it is a scan. A scanned PDF holds a photograph of each page and there are no letters in it to extract, only pixels arranged to look like them. Getting text from that needs OCR, which recognises the shapes as characters. The tool detects this case and says so rather than returning an empty box.

Why are some words run together or split?

A PDF stores text as positioned fragments rather than as sentences and some producers place every letter individually. The tool rebuilds lines from positions and inserts spaces where the gaps suggest them, which works for most documents. Fully justified text and unusual fonts occasionally fool it.

Will the layout be kept?

Lines are kept in reading order, top to bottom. Columns, tables and exact spacing aren’t reproduced, because plain text has no way to express them. For a table, copying from the PDF reader directly or converting to a spreadsheet with a dedicated tool, preserves more structure.

Does it read text in other languages?

Yes, any language the PDF stores as text, including accented Latin, Greek, Cyrillic, Chinese, Japanese and Arabic. The .txt download is saved as UTF-8, which every modern editor opens correctly.

Is the PDF uploaded?

No. The text is read in your browser by the same engine that displays PDFs and nothing is sent anywhere.

Guides

The other tools

All of them work the same way: the file is read in your browser and never uploaded.