Copied to clipboard
Processed on-device · pdf.js + Tesseract.js

Extract Text from Scanned PDFs

Turn scanned pages into text you can copy, search, edit, and export

Choose the document language and recognition mode, improve a rough scan when needed, then review and export the extracted text as TXT or JSON.

See the four steps
7 supported languages Scan enhancement filters Export TXT or JSON PDF processed in your browser
Built & maintained by Bhavin J. Sheth | Updated

Extract text from your PDF

Upload a scanned PDF, choose the language and recognition mode, then review and export the extracted text.

Drop a PDF here

or choose a file from your device

How OCR actually works

OCR turns pixels back into words

A scanned page is just a picture — there's no text underneath to select or search. OCR looks at the shapes of the letters and guesses which characters they are, one word at a time.

How the OCR engine processes a page

A photographed contract, an old scanned book, a receipt snapped on a phone — none of these carry a text layer, so there's nothing to copy or search until something reads the image and reconstructs the words. That's the entire job of optical character recognition.

pdf.js renders each page to a canvas; Tesseract.js analyzes that rendered image and returns recognized text with a confidence score. The PDF is processed in your browser and is not uploaded to an AllInOneTools server for OCR.

What each setting actually controls

Language

Loads the letter and word patterns Tesseract expects — matching the selected language to the document is one of the most important settings.

Accuracy mode

Choose Fast for quicker processing, or High Accuracy when you want a higher-resolution render before recognition.

Page selection

All pages, or a typed range — useful when only part of a long scan actually needs converting.

Enhancement filters

Grayscale and contrast preprocess the rendered page before OCR, which can help with some faded, noisy, or uneven scans.

Confidence score

A per-page number showing how sure Tesseract was — a quick signal for which pages to double-check by hand.

Settings that actually move accuracy

Language and accuracy mode matter far more than which filters you toggle.

1

Upload your PDF

Drag it onto the upload area, or click Select PDF. Thumbnails of every page appear right away.

Clear, high-resolution scans give the recognizer far more to work with.
2

Set language and accuracy

Use the control panel to configure recognition:

Match the language dropdown to what's actually printed on the page.
Choose Fast for speed, or High Accuracy for a sharper render before recognition.
Toggle Grayscale or Contrast for a rough or low-quality scan.
3

Extract

Click Extract Text. A progress bar tracks each page as Tesseract works through the document.

OCR runs in your browser; your PDF is not uploaded to an AllInOneTools server for processing.
4

Review and export

Check the confidence score per page, edit the text if needed, then copy it or download as TXT or JSON.

Per-page confidence TXT / JSON export

From a flat scan to text you can search

A scanned page contains image pixels instead of selectable text. OCR analyzes those pixels and produces editable text.

Scanned page Image only
Scanned PDF page with no selectable text before OCR

Nothing to copy, nothing to search — just a picture of text.

Recognized
Extracted text Sample output
Extracted searchable text after running OCR

Selectable, searchable text ready to copy, edit, or export.

Documents where the words are trapped in an image

Cases where information exists on the page but can't be searched, copied, or indexed until OCR runs.

Common situations for OCR text extraction from PDFs, by user type
SituationWhat gets extractedTypical user
Scanned research papersExtract quotable, searchable text from decades-old scanned journal articles.Researchers & academics
Legal document reviewMake scanned contracts and filings searchable for keyword review.Lawyers & paralegals
Receipt & invoice archivesPull line-item text from photographed receipts for expense tracking.Accountants & freelancers
Old family recordsConvert scanned letters and certificates into text that's easy to search later.Genealogists & families
Government formsExtract text from scanned public records for data entry or analysis.Public sector & researchers
Textbook digitizationTurn a scanned textbook chapter into text for accessibility tools or notes.Students & educators
Multilingual archivesRecognize non-English printed text using the matching language pack.Translators & archivists
Business card & label scansExtract short blocks of printed text from photographed cards or labels.Sales & admin teams
Accuracy factors

What actually improves recognition accuracy

Not every setting matters equally — here's what genuinely moves the needle, in rough order of impact.

Correct language selection

One of the most important settings. Using the wrong language model can significantly reduce recognition quality.

Source scan resolution

A 300 DPI scan of typed text gives Tesseract clean letter shapes to work with. A blurry phone photo or a low-DPI fax loses detail no setting can fully recover.

High accuracy mode

Renders each page at a higher scale before recognition, which helps on borderline scans — at the cost of noticeably longer processing time per page.

Enhancement filters

Grayscale and contrast help most on scans with color noise, faint print, or uneven lighting — they're a smaller boost, but a real one, on the right kind of image.

When the extracted text comes out wrong

What's actually happening, and the fastest way through it.

Symptom
Why it happens
Fix

Text is full of odd characters

The language setting doesn't match what's printed on the page.

Switch the language dropdown to match the document, then re-run extraction.

Confidence score is low

The source scan is blurry, faint, or low resolution.

Try High Accuracy mode, or turn on Grayscale and Contrast filters.

Processing takes a long time

High Accuracy mode and long documents both take longer, since each page runs recognition independently.

Use Fast mode for a first pass, or select specific pages instead of the whole document.

Some pages are missing from the output

The specific-pages range didn't include those page numbers.

Check the page range field, or switch to All Pages to be sure.

A password-protected PDF won't load

Password-protected PDFs may not be readable by the in-browser PDF renderer.

Remove the password protection first, then upload the PDF again.

Handwritten sections come out garbled

The recognizer is trained on printed text — handwriting, especially cursive, isn't reliably supported.

Expected — treat handwritten results as a rough draft to correct by hand.

Text is full of odd characters

Why: Wrong language selected.

Fix: Match the language, then re-run.

Low confidence score

Why: Blurry or low-resolution scan.

Fix: High Accuracy mode, or Grayscale/Contrast filters.

Processing is slow

Why: High Accuracy or a long document.

Fix: Use Fast mode, or select fewer pages.

Pages missing from output

Why: Page range didn't include them.

Fix: Check the range or switch to All Pages.

Password-protected PDF won't load

Why: Password-protected PDFs may not be readable by the in-browser PDF renderer.

Fix: Remove the password protection first, then upload the PDF again.

Handwriting comes out garbled

Why: Engine is built for printed text.

Fix: Expected — treat as a rough draft.

Getting clean text on the first pass

A practical order that helps avoid unnecessary reprocessing.

Get the language right before anything else

A mismatched language pack can significantly reduce recognition quality, so set the document language first.

Start with Fast mode to check the results

Run a quick pass first — only switch to High Accuracy if the confidence scores or text quality aren't good enough.

Use enhancement filters on rough scans only

Grayscale and contrast help faded or noisy scans, but can slightly hurt an already-clean, high-resolution image.

Check the confidence score per page

A lower score can help identify pages that deserve a closer manual review before you rely on the extracted text.

Edit directly in the results box

The extracted text is editable right there — fix obvious errors before copying or exporting, instead of after.

Export JSON when you need structure

JSON keeps each page's text and confidence score separate — useful for further processing instead of one flat text block.

PDF OCR questions

Answers about privacy, OCR accuracy, supported documents, handwriting, and processing.

Is it safe to upload my documents for OCR?

Your PDF is processed in your browser. The page does not upload the PDF to an AllInOneTools server for OCR.

How accurate is the text recognition?

Accuracy depends on scan quality, language, resolution, font, and layout. Clear printed scans generally produce better results, while blurry, skewed, noisy, or handwritten pages may need manual correction.

Can this tool read handwritten text?

The OCR engine is built for printed text. It may catch very neat, block-style handwriting, but cursive or messy handwriting will have low accuracy. Printed text gives the best results.

What file types work best?

Any PDF works, but image-based PDFs (scans) benefit most from OCR. If a PDF already has selectable text, you can usually copy it directly without running OCR at all.

What happens to my file after I'm done?

Nothing — it was never uploaded. The extracted text is generated in your browser; you copy or download it, and the original file stays untouched on your device.

Build notes

How this PDF OCR tool works under the hood

Because rendering, image preprocessing, and recognition all run client-side, I wrote up the implementation — the pdf.js-to-canvas pipeline, Tesseract.js worker setup, and the enhancement filters that clean up a scan before OCR.

Published on freeCodeCamp

How to Build a PDF OCR to Text Converter Using JavaScript

Covers browser-based OCR with pdf.js and Tesseract.js, and client-side document processing that keeps files private on the user's device.

Read the write-up on freeCodeCamp →

Written by Bhavin Sheth · Founder, AllInOneTools