PDF guides
Why Scanned PDFs Can't Convert to Text (No OCR)
Before you upload a PDF to any text-extraction tool, you can usually tell in a few seconds whether it will actually work — without needing to know anything about how PDFs are built. This guide starts with that practical, do-it-yourself check, then explains why PDF to Text detects and rejects true scanned PDFs rather than attempting a guess.
Ready to try the tool this guide describes?
Check your PDF yourself, before you upload it
- Open the PDF in any viewer — your browser, Preview, Adobe Reader — and try selecting text on the page with your cursor. If a sentence highlights, the page has real, extractable text.
- Try Ctrl+F (or Cmd+F) and search for a word you know appears on the page. A real hit confirms extractable text.
- If neither works — the page behaves like a single flat image, nothing highlights, and the search finds nothing — the PDF has no text layer at all.
Some scans already have hidden text — and those work fine
Not every scanned-looking PDF fails the check above. Many scanners and phone scanning apps — and some document-sharing or archival software — run OCR at scan time and embed the recognized words as an invisible text layer sitting behind the visible page image. Visually the page still looks like a photograph of paper, but the underlying PDF genuinely has real, extractable text, so it will pass the selection and search test above, and PDF to Text will convert it normally.
PDF to Text does not care whether a page looks like a scan; it only cares whether extractable text actually exists in the file.
Why true scanned PDFs are rejected instead of guessed at
For a PDF with no text layer, PDF to Text samples the first several pages and checks both how much extractable text they contain and how much of each page is covered by an embedded image. When a PDF has almost no extractable text across those sampled pages and is mostly covered by images, it is treated as scanned and the conversion is stopped before it starts, with a clear message shown instead of an empty or garbled .txt file.
What to do with a true scanned PDF
PDF to Text does not perform OCR. If you need text out of a genuinely scanned PDF with no text layer, you will need a dedicated OCR tool or service to recognize the text first, or re-export the document from its original source if you still have it.
Why Scanned PDFs Can't Convert to Text FAQ
- How can I tell if my PDF has real text before uploading it?
- Try selecting text on the page in any PDF viewer, or search for a word you know is there with Ctrl+F. If either works, the PDF has extractable text.
- My PDF looks like a scanned photo but has a text layer — will it work?
- Yes. If the page was OCR’d at scan time and has an invisible text layer behind the image, PDF to Text will find and extract that real text normally.
- Why does PDF to Text reject some PDFs instead of trying anyway?
- To avoid producing a misleading empty or garbled .txt file. A PDF with no extractable text is detected up front and rejected with a clear explanation instead.
- Does Looty Tools perform OCR?
- No. PDF to Text only extracts text that already exists in the PDF; it does not recognize text within scanned images.
- How does the tool detect a scanned PDF automatically?
- It samples the first several pages and checks the average extractable text per page alongside how much of each page is covered by an embedded image.
- What should I do if my PDF really is a scan with no text layer?
- Use a dedicated OCR tool or service to recognize the text first, or convert from the original source document if one is still available.
Related guides
- How to Convert PDF to TextTurn a text-based PDF into a plain .txt file in your browser, and see when plain text — not Word or HTML — is actually the right output format for the job.
- PDF to Text Line & Table HandlingThe specific mechanics behind PDF to Text's plain .txt output: pipe-joined table rows, the form-feed page-break marker, and why multi-column layouts can read out of order.
- Scanned PDFs and PDF to WordWhy scanned or image-only PDFs have no extractable text, how PDF to Word detects them up front, and why this tool does not perform OCR.