Scanned paper

Search scanned documents on your Mac, without changing a single file

Everything you scanned is on the disk and none of it is findable. That is not your filing being bad. It is that a scan contains no words — only a picture of them.

Loupe runs optical character recognition over scanned pages and photographs of documents on your own Mac, in twenty-eight languages, and indexes what it read. The scan itself is never modified — no text layer is written into it, and its date does not change. You search for what the page said, and the file comes up.

The thing nobody tells you about scans

Put a contract through a scanner and you get a PDF. Open it and you can read every word, so it is natural to assume the Mac can too. It cannot. What the scanner produced was a photograph of the page, saved inside a PDF wrapper. There is no text in the file — there are pixels arranged in the shape of text, which is a completely different thing to a computer and an identical thing to you.

The two-second test: open the file in Preview and try to drag-select a sentence. If nothing highlights, there is nothing there to find, and no amount of rebuilding an index will change that.

What optical character recognition actually does

OCR looks at the picture and works out which characters those shapes are, turning a photograph of the word “Invoice” into the word “Invoice”. It is a well-understood problem, and modern text recognition is good enough that printed pages come back close to perfect, including tables and columns.

Two things vary. Handwriting is much harder than print, and a page photographed at an angle in bad light gives up less than one laid flat. In practice this matters less than it sounds, because you are searching, not transcribing: you need enough of the page to come back for the supplier's name or the amount to be there, not a perfect copy.

What Loupe does with it

On the first run it opens every document in the folders you ticked. A PDF with real text in it is read directly. A PDF with no text — a scan — goes through OCR instead, and so does every photograph and screenshot on the disk. What comes back is stored in the index as ordinary text, so it is searchable by the same query box as everything else, and the row that comes back tells you which words earned the match.

It reads five pages of a scan — five pages that needed reading, rather than the first five in the file. Blank sides do not use up the budget, which matters more than it sounds: scanning double-sided produces a blank for every sheet, and counting them would mean stopping halfway through the document.

Five is a limit chosen on purpose. OCR is by far the most expensive thing the app does, and on a scanned document the things you would ever search for — who it is from, what it is about, the date, the total — are almost always at the front. Reading fifty pages of a contract to index the boilerplate would cost an hour of fan noise to find nothing you would have typed.

A document does not have to be all one thing

A report can be forty typed pages with five scanned exhibits stapled to the back, and that used to be the worst case here: the file carried plenty of text overall, so it was declared readable and the exhibits were never looked at. Not capped — skipped, and silently, because the file looked perfectly well indexed.

The question is now asked page by page. A page with its own text is taken as it is; a page without any is read with OCR wherever it sits in the file. And a page that has no text even after OCR — a photograph, a chart, a diagram — is described instead, so it is findable by what it shows rather than not at all.

Your files are not touched

A great many "make your scans searchable" tools work by rewriting the PDF and burying an invisible text layer inside it. That is a reasonable design, and it is not this one. Loupe keeps what it read in its own index, in a single file on your Mac. Your scan is byte-for-byte what it was, its modification date has not moved, and nothing has been re-saved by software you installed last week. Delete the index and nothing of yours is affected, because reading is all it ever did.

Twenty-eight languages for scans, every language for everything else

Text recognition covers twenty-eight languages, across Latin, Cyrillic, Arabic, Thai and the Chinese, Japanese and Korean scripts. That list is worth being precise about, because it has real holes: Greek, Hebrew and the Indic scripts are not in it. A scanned page in one of those has no text recovered from it, and stays findable by its filename and date the way it always was.

Everything above applies only to scans. A document that already contains text — a Word file, an exported PDF, a note — is indexed in any language and any script there is, with no model involved. Search for Küchenrechnung, 租赁合同 or ΑΝΑΚΑΙΝΙΣΗ and the right file comes back, including the Greek one that optical character recognition could never have read.

What it does not do

It does not correct your scans, deskew them, merge them, split them or file them. It does not add a text layer you can copy from in Preview — for that you want a PDF editor, and macOS itself will let you select text in an image with Live Text, one file at a time. This does one thing: it makes the pile findable.

Questions people ask about this

Why can’t my Mac find text in a scanned PDF?

Because the PDF does not contain any text. A scanner produces a photograph of the page and wraps it in a PDF, so the file holds pixels, not characters. Anything that searches by reading stored text has nothing to read. You can see this for yourself: open the file in Preview and try to select a line with the cursor. If nothing highlights, there is no text in it.

Does Loupe change my scanned files to make them searchable?

No, and this is deliberate. It reads the pages and stores what it read in its own index. Your PDF is not rewritten, no text layer is added to it, and its modification date does not move. The app has no code path that writes to your folders at all — it opens files, reads them, and closes them.

How much of a long scan does it read?

Five pages — but five pages that actually needed reading, not the first five in the file. Blank sides are skipped without using up the budget, which matters because double-sided scanning produces one for every page. Recognising text in an image is the most expensive thing the app does, and on a scanned document the supplier, the date, the amount and the subject are almost always at the front, so five is where the line sits.

Which languages does the OCR handle?

Twenty-eight, across Latin, Cyrillic, Arabic, Thai and the Chinese, Japanese and Korean scripts. Greek, Hebrew and the Indic scripts are not among them — a scan in one of those yields no text, and stays findable by its name and date only. Documents that already contain text are a different matter: those are searchable in any language at all.

Does the OCR happen on a server?

No. It runs on your Mac, using the text recognition built into macOS and models that live on the machine. Your scans are never uploaded, and neither are their filenames. Search works with the network off.

What about a photo I took of a receipt instead of scanning it?

Same treatment. A photograph of a document is read for its text exactly as a scan is, which is usually more useful than scanning, because everybody has a phone and almost nobody has a scanner.

Point it at the folder full of scans.

It reads them once, in the background, on your own machine. Then search for the supplier you can picture but cannot name.

Try it free7 days, no card neededSee pricing
  • Apple silicon · macOS 26 or later
  • No account
  • Nothing leaves your Mac