Illustration of a text document page and a picture-only page each compared through a different lens icon

PDF guides

Comparing Scanned PDFs vs. Text PDFs: What's Different

Not every PDF has real text to compare. Here’s exactly what happens when Compare PDF encounters a scanned or image-only page, and why the result is still useful.

Ready to try the tool this guide describes?

Why scanned pages can't get a text diff

A scanned or photographed PDF page is really just a picture of a document — there’s no underlying text to extract at all, so there’s nothing for a line-by-line text comparison to work from.

How the tool tells the two apart

Each page is checked for how much real, extractable text it actually contains. A page with little to no extractable text is treated as scanned; a page with a substantial amount of real text is treated as a text page. This is decided per page, not per document, so a PDF that mixes typed pages with a scanned signature page is handled correctly.

What you still get for scanned pages

Scanned pages fall back to a visual comparison of the rendered page image — still matched by content (not just position) across the two files, and still reported as unchanged, changed, added, or removed. You lose the line-level "what changed" detail, but you still get accurate page-level change detection.

Mixed documents

In a document with both text and scanned pages, each page uses whichever signal actually applies to it — text pages get the full text-and-visual check, scanned pages get the visual-only check — rather than the whole document falling back to the weaker signal just because one page is scanned.

Comparing Scanned vs. Text PDFs FAQ

Will comparing two scanned PDFs work at all?
Yes — they’re compared visually instead of by text, and you still get accurate Changed/added/removed results per page.
Can I get a line-level text diff for a scanned page?
No — there’s no extractable text on a scanned page for a line-level diff to work from; the visual comparison is the available signal.
How does the tool decide a page is "scanned" versus "text"?
By how much real, extractable text that specific page actually contains — decided independently for every page, not for the whole document at once.
What if one PDF is scanned and the other has real text?
Each page pair still gets matched by whatever signal applies — a scanned page is compared visually even against a text-page counterpart.
Does a scanned page ever get a false "Changed" from rendering noise alone?
The visual comparison uses a tolerance for small rendering/anti-aliasing differences so trivial noise alone shouldn’t trigger a false Changed result.

Related guides

Open the tool

Jump into Compare PDF when you are ready to process your files.

← Back to all guides

More from Looty

Explore Looty’s Ecosystem

Discover more ways Looty can help you learn, organize, create, and make an impact.