Back to tools

OCR PDF

Convert scanned PDFs into searchable and selectable documents.

Drop your PDF here or click to browse

PDF or image files supported

OCR language:

OCR PDF — turn scanned PDFs into searchable text

Extract text from image-only or scanned PDFs using open-source Tesseract OCR.

OCR PDF runs optical character recognition (Tesseract 5) on every page of your PDF and returns the extracted text. Use it when a PDF was created by scanning paper and the pages are images — a normal 'copy text' or PDF-to-Word won't work on those without OCR first.

Language packs installed on our server: English, Hindi, Spanish and French. Accuracy depends on scan quality — a clean 300-DPI black-and-white scan gives near-perfect results; a phone photograph of a crumpled page gives noticeably worse results.

How to use it

  1. 1Upload the scanned PDF.
  2. 2Choose the language of the document (English is the default).
  3. 3Press OCR PDF and wait — a page or two takes a couple of seconds, hundreds of pages can take longer.
  4. 4Copy the extracted text or download it. Feed it into PDF to Word if you want an editable version.

Real-world use cases

  • Search a scanned book by keyword.
  • Extract quotes from a photograph of a historical document.
  • Prepare an image-only PDF for translation with the Translate tool.
  • Feed the extracted text into an accessibility screen reader.

Supported formats & limits

Input: PDF. Supported OCR languages: English (eng), Hindi (hin), Spanish (spa), French (fra). Unknown language codes fall back to English.

Privacy

Your uploaded file is written to a temporary folder on our server, processed, and deleted by the operating system as soon as the request finishes — typically within seconds. Nothing is stored in our database. Traffic is encrypted over HTTPS.

Full details in the Privacy Policy.

Frequently asked questions

How accurate is the OCR?

Clean printed text at 300 DPI: 95-99% accurate. Phone photos, handwriting, mixed languages: significantly lower. Always spot-check the output.

Which languages are supported?

English, Hindi, Spanish and French are installed on our server. Extending the list requires adding the corresponding Tesseract language pack.

Can it OCR handwriting?

Tesseract is designed for printed text. Handwriting works only for very clear samples and is unreliable — dedicated handwriting-recognition tools work better.

Will the output preserve columns and tables?

OCR produces plain text in reading order. Multi-column layouts are usually flattened. For preserving tables, try PDF to Excel on the OCR-ed PDF.

Related PDF tools