OCR a PDF

Turn a scanned PDF into one you can search. Runs in this browser tab, with no account, no upload and no page limit.

Why a scanned PDF needs OCR

A scanned PDF is a stack of photographs. Your eyes read it fine, but as far as the computer is concerned there is nothing on those pages but coloured dots. Ctrl+F finds nothing. You cannot copy a line out of it. Nothing indexes it, so it will not come back in a search later.

OCR — optical character recognition — is the step that changes that. It looks at the shapes on each page, decides which letters they are, and writes those letters into the file as an invisible layer underneath the image. Run OCR on a PDF and the pages look exactly the same; the difference is that now there is text under them.

The moment that matters is usually months later. A contract you scanned, an invoice, a receipt, a page from a book — you remember one phrase and nothing else. A searchable PDF gives that phrase back to you. A stack of photographs does not.

This PDF OCR tool runs entirely in your browser. The file is not uploaded, there is no account, and there is no cap on how many pages you can run. The page ranking above us for this keyword processes your file free and then asks you to sign in before it will give it back.

Upload toolsTheir server?This toolYour deviceNothing leaves your device, so there is no copy to delete.
Other converters upload your file to a server. This one draws it where it already is.

What is OCR on a PDF?

OCR on a PDF is character recognition applied to the page images inside it, with the recognised words written back into the same file as an invisible text layer — so the document looks unchanged but becomes searchable, selectable and copyable.

Two kinds of PDF look identical on screen. One was exported from a word processor and carries real text; the other was produced by a scanner or a phone camera and carries pictures of text. Only the second kind needs OCR.

There is a quick way to tell them apart without any tool: try to select a line with your mouse. If a blue highlight follows the words, the text is already there. If the whole page highlights as one block, or nothing happens at all, you have a scanned PDF.

This tool does that check for you before it runs anything. If the file already has a text layer, it says so and suggests you extract the text instead, because OCR would replace something exact with something approximate — and you would have no way to notice afterwards, since the page looks the same either way.

The text OCR produces goes underneath the page image, not on top of it. That is why a searchable PDF still looks like a scan: you are always looking at the original picture. The invisible layer is only there for your search box, your copy command, and whatever indexes the file.

It is worth being clear about what that layer is. It is a guess — a very good one on clean print, a poor one on a crumpled receipt photographed at an angle. The picture you see is always right. The text underneath it may not be.

Image onlyInvoiceNo. 4471Total1,240.00+ text layer
OCR does not change how the page looks. It adds a layer of text underneath the image, which is what makes the PDF searchable.

How to OCR a PDF in four steps

Everything happens in this browser tab. The file does not leave your device, and you will not be asked to create an account before you can download the result.

Lorem ipsumdolor sit ametconsecteturHas textSkip OCRLorem ipsumSome pagesSkip thoseImage onlyRun OCR
Not every PDF needs OCR. This tool checks before it runs anything, because re-reading text that is already exact only makes it worse.
1

Drop your scanned PDF in

Drag the file onto the box above, or click it to browse. Files up to 100 MB are fine. Nothing is sent anywhere — the file is opened by this page, in your browser.

2

Read the verdict

The tool checks every page for an existing text layer before it does any work. It will tell you whether your PDF needs OCR at all, needs it on some pages, or does not need it. If only some pages are scans, the rest are left untouched by default.

3

Pick the document's language

Accuracy depends on this more than on anything else you can control. Choose the language the document is written in, not your interface language. Language data is downloaded once — about 5 MB for English, 1.7 MB for Hindi — and then cached. Selecting several languages at once works, but slows every page down.

4

Run it and check the result

Pages are processed one at a time, with a preview of the page being read, so you can stop early if you dropped in the wrong file. When it finishes, try selecting some text in the preview. Then download the searchable PDF — or the plain text, if that is all you wanted.

Only need the words? If you want the text and not the PDF, and your file already has a text layer, there is a faster route that does not involve guessing at all: extract the text directly from the PDF.

What this OCR PDF tool does

It adds a text layer and changes nothing else. Your original pages are kept exactly as they are — same images, same resolution, same page boxes. The recognised words are drawn over them at zero opacity, so they can be selected and searched but never seen. Pages that already had text are not rewritten at all.

It renders pages at 150 DPI, which is a measured choice rather than a default. On a test document, accuracy at 100, 150, 200 and 300 DPI was the same within half a percentage point, while 300 DPI took four times as long. Below 100 DPI there is a cliff: at 72 DPI accuracy fell from 95.6% to 79.3%. Scanning at 600 DPI will not make OCR more accurate. It will only make it slower.

It handles more than 100 languages, including scripts most converters skip. One of the better-known online OCR tools lists 19 languages and not one of them is a South Asian script. Hindi, Tamil and Bengali are in the list here, near the top, because that is where most people searching for this actually are.

It counts what it could not do. If a page comes back with no text — a blank sheet, a photograph, handwriting — it stays in the file untouched and gets reported. Nothing is silently dropped.

It never asks you to sign in. You can run a 400-page document and then run another one. There is no daily limit, because there is no server doing the work and therefore nothing for us to ration.

People use PDF OCR for:

  • Making scanned contracts and agreements searchable before filing them
  • Finding one line in a book chapter that was photographed page by page
  • Turning a folder of scanned invoices into something you can grep
  • Getting text out of a fax or a photocopy that arrived as a PDF
  • Making old scanned records findable by search instead of by memory
  • Preparing scanned documents so a search index or another tool can read them
  • Copying a paragraph out of a scan instead of retyping it

Works on any device

There is nothing to install. The OCR runs in the browser you already have, which means the same tool works everywhere and your file stays on the device it started on.

§ 01

Windows

Chrome, Edge or Firefox. Nothing to download, and no Acrobat subscription needed to make a scan searchable.

§ 02

Mac

Safari or Chrome. Preview can show you a scanned PDF but it cannot add a text layer to one.

§ 03

iPhone and iPad

Works in Safari. Handy when the scan is a photo you just took — though a phone photo is exactly the case where straightening the page matters most.

§ 04

Android

Works in Chrome. Long documents will take a while and will keep the phone busy, so plug it in first.

§ 05

Chromebook

A good fit, because everything runs locally in the browser rather than needing an installed application.

§ 06

Linux

Any modern browser. No packages, no command line, and no queue.

What OCR gets right, and what it gets wrong

Every OCR tool makes mistakes, and almost none of them will tell you what those mistakes look like. Here is a real result: a clean, computer-rendered invoice, no noise, no skew, 95.6% character accuracy overall. Of twelve words sampled from it, five could be found by searching.

In the fileWhat OCR readNorthwindNorthwindSubtotalSubtotal1741.501741.502089.802089. 80INV-2026INV2026ReplacementRepl acenentPayment isPaymentis5 of 12 words searchable
From a clean test invoice at 95.6% character accuracy. The company name came through; the total did not, because OCR put a space inside the number.
In the documentWhat OCR readWhy
NorthwindNorthwindOrdinary prose in a proportional font is the easy case.
2089.802089. 80A space appeared inside the number, so searching for the total fails.
INV-2026-04471INV2026-04471One hyphen was dropped. Reference numbers are unusually fragile.
ReplacementRepl acenentAn m read as an n, plus a space that is not there.
Payment isPaymentisThe opposite failure: a real space was removed.
1741.501741.50Same table, same font as the total above — and this one was fine. The errors are not predictable.

The thing that actually ruins a scan

It is not compression, and it is not resolution. A JPEG at quality 60 scored within a fraction of a point of a lossless render. What hurt was skew: a page tilted by two degrees was fine, but at seven degrees accuracy dropped by nine and a half points and confidence fell from 91 to 74. If you are scanning something, straighten it. Do not bother raising the DPI.

Searchable PDF, plain text, or a Word file?

A searchable PDF is what this page makes: the scan you had, plus an invisible text layer, still a PDF. Plain text is the words with the layout thrown away — useful when you want to paste them somewhere, and offered here as a second download. An editable Word document is a different job that involves rebuilding the layout, and this tool does not do it. If your PDF already has text and you only want the words, skip OCR entirely and extract them directly.

Your scan never leaves this browser

The PDF you drop here is opened by this page, read by your own processor, and written back out by your own browser. It is not uploaded, not queued on a server, and not held anywhere for a retention period, because there is no server involved in the work at all.

This matters more for OCR than for most conversions. The documents people need to make searchable are contracts, medical records, bank statements and identity papers — the scans that exist precisely because the original was worth keeping. Every major online OCR service processes those on its own machines and deletes them on a schedule, typically after one to eight hours.

It also means there is nothing to ration. Services that do the work on their own hardware have to limit free use, which is why they cap you at a couple of tasks a day or ten pages a file. Here the work happens on your device, so there is no limit to enforce and no sign-in to collect.

§ 01

Nothing is uploaded

The file is read locally. No copy of it is created on any server.

§ 02

No account, ever

You are not asked to sign in to download the result, which is the usual charge for free OCR elsewhere.

§ 03

No page or file limit

Run a 400-page scan, then run another one. There is no daily quota.

§ 04

Language data only

The one thing downloaded is the recognition data for the language you pick. Your document is never part of that request.

Frequently asked questions

Is this PDF OCR tool really free, and is there a page limit?
It is free with no page limit and no daily quota. The OCR runs on your own device rather than on a server, so there is no per-file cost to us and nothing to ration. Big documents are limited only by your own patience and battery.
Do I need to sign in to download the result?
No. You can drop a file in, run OCR and download the searchable PDF without an account or an email address. This is worth checking on any OCR site before you upload something: it is common to be asked for an account only at the download step, after the work is done.
How do I know if my PDF even needs OCR?
Try to select a line of text with your mouse. If a highlight follows the words, the file already has a text layer and OCR would only make it worse. This tool checks automatically when you add a file and tells you which of the three cases you are in before anything runs.
How accurate is the OCR?
Accurate enough to be useful and not accurate enough to trust blindly. On a clean, computer-generated test invoice, character accuracy was 95.6% — but only 5 of 12 sampled words could actually be found by searching, because a space landed inside a number or a hyphen went missing. Clear printed text does well. Anything creased, faint or tilted does worse.
Which languages are supported?
Over 100, including Hindi, Tamil, Bengali, Thai, Vietnamese and Indonesian alongside the usual European languages. Pick the language the document is written in, not the language of this page — it affects accuracy more than any other setting. Note that scripts which reshape their characters, such as Devanagari, can lose the occasional word when it is written into the PDF text layer even though it was recognised correctly.
Will it read handwriting?
Assume not. This engine is built for printed characters, and handwriting is a different problem that it was not trained for. A handwritten page will usually come back empty or close to it. If that happens, the page is left in your file exactly as it was and reported in the summary rather than quietly dropped.
Will the PDF look different afterwards?
No, and that is the point. Your page images are kept exactly as they were and the recognised text is drawn over them at zero opacity. You keep looking at the original scan; the text layer is only there for your search box and your copy command.
Does a higher DPI scan give better OCR results?
Usually not. In testing, accuracy at 100, 150, 200 and 300 DPI was identical within half a percentage point, while 300 DPI took four times as long to process. There is a floor — at 72 DPI accuracy collapsed to 79.3% — but above about 150 DPI you are paying in time for nothing. Straightening the page helps far more.
How long does it take?
Roughly a second per page on a normal laptop for a standard A4 page, so a 50-page document lands around a minute. Phones are slower. Selecting several languages at once multiplies the work, so pick only the ones that are actually in the document.
Is my file uploaded anywhere?
No. The PDF is opened and processed inside this browser tab and never sent to a server. The only network request the tool makes is for the recognition data of the language you selected, which is the same file for everyone and contains nothing of yours.
What is the difference between this and PDF to Text?
This page makes a searchable PDF out of a scan by guessing at the characters in the images. PDF to Text pulls out text that is already in the file, exactly, with no guessing. If your PDF has a text layer, that is the tool you want — it is faster and it cannot make the kind of mistakes OCR makes.
Can I get an editable Word document instead?
Not from this tool. Turning a scan into a Word file means rebuilding the layout — columns, tables, headings — and that is a different job with its own failure modes. What you get here is the original PDF made searchable, plus a plain text download if you want the words on their own.

Make your scanned PDF searchable

Free, unlimited, and it runs in this tab. No upload, no sign-in, and it will tell you if your file does not need OCR at all.

OCR a PDF