Extract Text from Scanned PDFs
Turn scanned pages into text you can copy, search, edit, and exportChoose the document language and recognition mode, improve a rough scan when needed, then review and export the extracted text as TXT or JSON.
Extract text from your PDF
Upload a scanned PDF, choose the language and recognition mode, then review and export the extracted text.
Drop a PDF here
or choose a file from your device
Document preview (0 pages)
OCR turns pixels back into words
A scanned page is just a picture — there's no text underneath to select or search. OCR looks at the shapes of the letters and guesses which characters they are, one word at a time.
How the OCR engine processes a page
A photographed contract, an old scanned book, a receipt snapped on a phone — none of these carry a text layer, so there's nothing to copy or search until something reads the image and reconstructs the words. That's the entire job of optical character recognition.
pdf.js renders each page to a canvas; Tesseract.js analyzes that rendered image and returns recognized text with a confidence score. The PDF is processed in your browser and is not uploaded to an AllInOneTools server for OCR.
What each setting actually controls
Language
Loads the letter and word patterns Tesseract expects — matching the selected language to the document is one of the most important settings.
Accuracy mode
Choose Fast for quicker processing, or High Accuracy when you want a higher-resolution render before recognition.
Page selection
All pages, or a typed range — useful when only part of a long scan actually needs converting.
Enhancement filters
Grayscale and contrast preprocess the rendered page before OCR, which can help with some faded, noisy, or uneven scans.
Confidence score
A per-page number showing how sure Tesseract was — a quick signal for which pages to double-check by hand.
Settings that actually move accuracy
Language and accuracy mode matter far more than which filters you toggle.
Upload your PDF
Drag it onto the upload area, or click Select PDF. Thumbnails of every page appear right away.
Set language and accuracy
Use the control panel to configure recognition:
Extract
Click Extract Text. A progress bar tracks each page as Tesseract works through the document.
Review and export
Check the confidence score per page, edit the text if needed, then copy it or download as TXT or JSON.
From a flat scan to text you can search
A scanned page contains image pixels instead of selectable text. OCR analyzes those pixels and produces editable text.
Nothing to copy, nothing to search — just a picture of text.
Selectable, searchable text ready to copy, edit, or export.
Documents where the words are trapped in an image
Cases where information exists on the page but can't be searched, copied, or indexed until OCR runs.
| Situation | What gets extracted | Typical user |
|---|---|---|
| Scanned research papers | Extract quotable, searchable text from decades-old scanned journal articles. | Researchers & academics |
| Legal document review | Make scanned contracts and filings searchable for keyword review. | Lawyers & paralegals |
| Receipt & invoice archives | Pull line-item text from photographed receipts for expense tracking. | Accountants & freelancers |
| Old family records | Convert scanned letters and certificates into text that's easy to search later. | Genealogists & families |
| Government forms | Extract text from scanned public records for data entry or analysis. | Public sector & researchers |
| Textbook digitization | Turn a scanned textbook chapter into text for accessibility tools or notes. | Students & educators |
| Multilingual archives | Recognize non-English printed text using the matching language pack. | Translators & archivists |
| Business card & label scans | Extract short blocks of printed text from photographed cards or labels. | Sales & admin teams |
What actually improves recognition accuracy
Not every setting matters equally — here's what genuinely moves the needle, in rough order of impact.
Correct language selection
One of the most important settings. Using the wrong language model can significantly reduce recognition quality.
Source scan resolution
A 300 DPI scan of typed text gives Tesseract clean letter shapes to work with. A blurry phone photo or a low-DPI fax loses detail no setting can fully recover.
High accuracy mode
Renders each page at a higher scale before recognition, which helps on borderline scans — at the cost of noticeably longer processing time per page.
Enhancement filters
Grayscale and contrast help most on scans with color noise, faint print, or uneven lighting — they're a smaller boost, but a real one, on the right kind of image.
When the extracted text comes out wrong
What's actually happening, and the fastest way through it.
Text is full of odd characters
The language setting doesn't match what's printed on the page.
Switch the language dropdown to match the document, then re-run extraction.
Confidence score is low
The source scan is blurry, faint, or low resolution.
Try High Accuracy mode, or turn on Grayscale and Contrast filters.
Processing takes a long time
High Accuracy mode and long documents both take longer, since each page runs recognition independently.
Use Fast mode for a first pass, or select specific pages instead of the whole document.
Some pages are missing from the output
The specific-pages range didn't include those page numbers.
Check the page range field, or switch to All Pages to be sure.
A password-protected PDF won't load
Password-protected PDFs may not be readable by the in-browser PDF renderer.
Remove the password protection first, then upload the PDF again.
Handwritten sections come out garbled
The recognizer is trained on printed text — handwriting, especially cursive, isn't reliably supported.
Expected — treat handwritten results as a rough draft to correct by hand.
Text is full of odd characters
Why: Wrong language selected.
Fix: Match the language, then re-run.
Low confidence score
Why: Blurry or low-resolution scan.
Fix: High Accuracy mode, or Grayscale/Contrast filters.
Processing is slow
Why: High Accuracy or a long document.
Fix: Use Fast mode, or select fewer pages.
Pages missing from output
Why: Page range didn't include them.
Fix: Check the range or switch to All Pages.
Password-protected PDF won't load
Why: Password-protected PDFs may not be readable by the in-browser PDF renderer.
Fix: Remove the password protection first, then upload the PDF again.
Handwriting comes out garbled
Why: Engine is built for printed text.
Fix: Expected — treat as a rough draft.
Getting clean text on the first pass
A practical order that helps avoid unnecessary reprocessing.
Get the language right before anything else
A mismatched language pack can significantly reduce recognition quality, so set the document language first.
Start with Fast mode to check the results
Run a quick pass first — only switch to High Accuracy if the confidence scores or text quality aren't good enough.
Use enhancement filters on rough scans only
Grayscale and contrast help faded or noisy scans, but can slightly hurt an already-clean, high-resolution image.
Check the confidence score per page
A lower score can help identify pages that deserve a closer manual review before you rely on the extracted text.
Edit directly in the results box
The extracted text is editable right there — fix obvious errors before copying or exporting, instead of after.
Export JSON when you need structure
JSON keeps each page's text and confidence score separate — useful for further processing instead of one flat text block.
PDF OCR questions
Answers about privacy, OCR accuracy, supported documents, handwriting, and processing.
Is it safe to upload my documents for OCR?
Your PDF is processed in your browser. The page does not upload the PDF to an AllInOneTools server for OCR.
How accurate is the text recognition?
Accuracy depends on scan quality, language, resolution, font, and layout. Clear printed scans generally produce better results, while blurry, skewed, noisy, or handwritten pages may need manual correction.
Can this tool read handwritten text?
The OCR engine is built for printed text. It may catch very neat, block-style handwriting, but cursive or messy handwriting will have low accuracy. Printed text gives the best results.
What file types work best?
Any PDF works, but image-based PDFs (scans) benefit most from OCR. If a PDF already has selectable text, you can usually copy it directly without running OCR at all.
What happens to my file after I'm done?
Nothing — it was never uploaded. The extracted text is generated in your browser; you copy or download it, and the original file stays untouched on your device.
How this PDF OCR tool works under the hood
Because rendering, image preprocessing, and recognition all run client-side, I wrote up the implementation — the pdf.js-to-canvas pipeline, Tesseract.js worker setup, and the enhancement filters that clean up a scan before OCR.
How to Build a PDF OCR to Text Converter Using JavaScript
Covers browser-based OCR with pdf.js and Tesseract.js, and client-side document processing that keeps files private on the user's device.