Why can’t I copy text from a PDF?
Find out whether your PDF contains selectable text or scanned images and choose the right way to extract readable words from it.
A PDF that looks full of words may contain nothing you can select. The first question is whether those letters are stored as text or as an image. A different extraction method can help but repeatedly copying the same page won’t change what it contains.
Try one sentence first
Open the PDF in a reader and drag across a short sentence. Copy it into a plain text editor. If the result is readable then start with PDF to Text. It extracts an existing text layer instead of guessing letters from their shapes.
If the whole page behaves like one picture then try Image to Text. It accepts scanned PDFs and uses optical character recognition to identify words. A scan with an existing OCR layer may already allow selection so appearance alone isn’t enough to classify it.
If copying is explicitly disabled or the reader requests a password then ask the document owner for an accessible copy. Don’t assume every failure means the page is a scan.
Choose the route that matches the result
| What you observe | Useful next step |
|---|---|
| Text copies correctly | Extract the existing text |
| Nothing can be selected | Try OCR on one representative page |
| Some pages copy but others don’t | Handle the scanned pages separately |
| Copied characters are garbled | Compare another reader and the source document |
| Words are correct but columns are mixed | Review the reading order manually |
For example a ten-page report might contain nine exported pages followed by a photographed receipt. Extract the report’s text normally. Recognise the receipt separately and label the output so you know where it came from.
Check the output before you rely on it
Read names and reference numbers against the original. Pay particular attention to zero versus the letter O and one versus lowercase l. A recogniser can return a plausible word that wasn’t actually printed.
For a two-column page confirm that the first column finishes before the second begins. A PDF positions characters on a page but the intended reading order isn’t always straightforward. Correct words in the wrong order can still change the meaning of a sentence.
Keep page references when extracting several pages. A note such as “receipt on page 10” makes it much easier to resolve a doubtful amount later. Keep the original PDF unchanged while you work.
Extracting text isn’t the same as making a searchable PDF
Utility Mule’s text tools produce text you can copy or save. Don’t expect that output to add a searchable layer back into the original PDF. If you need that particular deliverable then use a workflow that explicitly saves a searchable PDF. Adobe describes that distinction in its OCR instructions.
If recognition is poor then return to the scan. A straighter page with readable small print is more useful than repeatedly processing the same blurred image.
Common questions
Does every PDF contain text?
No. A PDF can contain a photograph of a page with no text characters behind it. It can also mix scanned pages with pages exported from a word processor. Check more than one page before choosing how to extract the text.
Will OCR preserve my table layout?
Not reliably. Recognition can produce readable words while losing column relationships. Compare each row with the original before using the output as a spreadsheet or treating adjacent numbers as related values.
Should I photograph the screen instead?
Use the original PDF when you have it. Photographing a screen introduces reflections and blur that make recognition harder. If a single page needs attention then extract that page rather than degrading the entire document.