PDF to Word Converter — editable .docx with OCR and dynamic page geometry in your browser
Converts digital and scanned PDFs, preserving page dimensions (A4, Letter, Landscape), bold, italic, headings, headers, and footers — before your file downloads.
Page dimensions, tables, and images are preserved. Word page sizes and orientations match your PDF automatically. Review your converted DOCX in Microsoft Word or LibreOffice for final document sign-off.
The conversion console
Upload a PDF, customize document setup, then convert.
Browser-based processing
PDF parsing, page dimension extraction, OCR, and Word document generation run securely inside your browser.
PDF to Word
Drop your PDF here
or browse files
PDF page
Extracted text
Extracted text preview
The exported DOCX file retains matched page sizes, margins, font styles, and header/footer configurations.
Fix any OCR mistakes here. Manual edits replace paragraph runs with adjusted plain text.
Export settings
Full formatting: page sizing, headings, bold, italic, tables, images, headers/footers.
Text extraction and formatting capabilities
Designed to extract and structure PDF text reliably while being transparent about current document limitations.
What it supports
- Runs OCR (Tesseract) on scanned pages, with automatic detection of which pages need it.
- Detects bold, italic, and heading-sized text from the PDF's own fonts and carries them into real Word formatting.
- Builds a navigable Word outline from detected headings.
- Prompts for a password on protected PDFs instead of failing.
- Lets you convert a page range, add a header/footer/title page, and choose font, spacing, and margins.
- Detects tables, underlined text, and images and turns them into editable Word tables, underline formatting, and embedded pictures.
Current limitations
- Embedded images are placed in reading order, not at pixel-exact positions due to flowing layout differences.
- Underline detection granularity depends on how text runs are grouped in the source PDF.
- Paragraphs you edit in the correction preview lose bold/italic/heading formatting as plain text is substituted. Editing OCR text may simplify some original formatting.
- Only one PDF at a time is processed in the current single-file flow.
From PDF to editable Word document
Upload, configure, then convert — password-protected files just add one extra step.
Upload your PDF
Drag it onto the console, or click Select PDF File. If it's password protected, you'll be asked to unlock it first.
Set your options
Configure exactly what you need before converting:
Convert
Click Convert to Word. Enable side-by-side view first if you want to watch each page's extracted text as it's processed.
Review and download
Check the preview, then choose a file name and format — .docx, .md, or .txt.
What this converter is suited for
And where document complexity might require additional review.
| Use case | Why | Fit |
|---|---|---|
| Editing a scanned contract | OCR turns scanned pages into editable text, with bold/italic and headings preserved. | Strong fit |
| Repurposing a report's text | Headings become a real, navigable Word outline; body text keeps its formatting. | Strong fit |
| Converting just a few pages | Page range selection avoids converting an entire long document. | Strong fit |
| A document full of data tables | Enable "Detect tables" to convert detected aligned columns into editable Word tables. | Good fit |
| A brochure with photos or diagrams | Enable "Extract and embed images" to bring pictures into the output, placed in reading order. | Good fit |
| Converting dozens of PDFs at once | Only one file is processed at a time currently. | Not yet a fit |
What's included in this version
Everything marked "supported" is built and functional. Everything marked "planned" is kept separate for future releases.
Password-protected PDFs Supported
You're prompted for the password and conversion proceeds normally once verified, instead of failing outright.
Page range selection Supported
Enter a range like "1-3, 7, 10-12" to convert only those specific pages instead of the entire document.
Header on every page Supported
Custom text, file name, or today's date placed at the top of every page via OOXML headers.
Font, spacing & margin options Supported
Pick a font family, base size, and compact/normal/wide presets for line spacing and document margins.
Title page Supported
An optional cover page with document title and date ahead of the main content.
Italic preservation Supported
Detected from font information and carried into the Word file as real italic formatting.
Heading detection Supported
Distinct larger font sizes are mapped to Word Heading 1/2/3 styles for a navigable outline.
Side-by-side view Supported
Watch rendered PDF pages next to extracted text as conversion progresses page by page.
Dark mode for the console Supported
A toggle for the converter interface, persisting your preference locally.
.txt and .md export Supported
Export as Markdown syntax or plain text alongside standard .docx files.
Batch conversion + ZIP Planned
Requires a dedicated multi-file queue with per-file progress tracking in future updates.
Table detection Supported
Aligned columns of text are parsed and converted into Word tables based on spatial positioning.
Image extraction & embedding Supported
Images are extracted and embedded in Word files in reading order per page.
OCR image preprocessing Supported
Optional noise reduction and contrast binarization run before OCR processing.
Underline preservation Supported
Detects thin drawn lines in PDFs and correlates them with nearby text runs.
Editable OCR-correction preview Supported
Review and edit extracted text prior to final document generation.
Troubleshooting common conversion issues
Understand why issues happen and how to resolve them quickly.
"No text could be extracted"
The PDF is scanned and OCR is set to Off, or the OCR language doesn't match the document's actual language.
Switch OCR mode to Auto or Always, and check the language setting.
A table looks like scrambled text
The "Detect tables" option wasn't enabled for this conversion, so columns were extracted as plain wrapped text.
Rebuild the table manually, or re-convert with "Detect tables" turned on.
Some text isn't bold/italic anymore
Detection relies on font metadata in the PDF. Scans or flattened files often lack explicit style names.
Reapply formatting manually in Word where needed; this is a source-file limitation.
"Incorrect password" on a password I'm sure is right
PDF passwords are case-sensitive. Some PDFs also have a separate permissions password, which is different from the password required to open the file.
Double check capitalization and ensure you are entering the user/open password.
OCR is slow
OCR runs entirely in your browser and takes processing time per page, especially during initial language downloads.
This is expected behavior — set OCR mode to Auto so it only runs on scanned pages.
"No text could be extracted"
Why: OCR off, or scanned PDF.
Fix: Set OCR to Auto/Always, check language.
Table looks scrambled
Why: "Detect tables" wasn't enabled.
Fix: Re-convert with it turned on.
Bold/italic missing
Why: Source PDF lacks font metadata.
Fix: Reapply manually in Word.
"Incorrect password"
Why: Case-sensitive or permissions password used.
Fix: Confirm exact open password.
OCR is slow
Why: Runs locally in browser.
Fix: Normal — use Auto mode.
Best practices for clean conversions
Recommended approaches for achieving optimal conversion results.
Set the OCR language correctly
Tesseract accuracy relies on matching the language setting to the document's actual language.
Use a page range for long documents
Converting specific pages is faster and easier to review when you only need a section of a large PDF.
Turn on side-by-side view for scans
Comparing extracted text directly against the rendered page helps spot OCR discrepancies instantly.
Check heading outlines in Word
Open Word's Navigation pane after converting to verify that headings align with document structure.
Keep original PDFs as backup
Retain source files until complex tables or specialized layouts are fully verified in the DOCX output.
Export as Markdown for wikis
When preparing documentation for Markdown-based platforms, the .md export retains formatting directly.
PDF to Word FAQs
Find answers about scanned PDFs, formatting, page ranges, passwords, privacy, and converted Word documents.
Can I convert scanned PDFs to editable Word documents?
Yes. This converter uses Tesseract.js OCR in your browser to recognize text from scanned pages. In Auto mode, pages with little or no extractable text are sent through OCR. Results depend on scan quality, document layout, and the selected language, so review the converted document before using it.
Does this preserve formatting like bold, italic, and headings?
Bold and italic are detected from the PDF's own font information and carried into the Word file as real character formatting. Larger, distinct font sizes are detected as headings and mapped to Word's built-in Heading 1/2/3 styles, which gives the document a real, navigable outline in Word — not just bigger text.
Does this preserve underlined text?
Yes. Underlines in PDFs are usually drawn as thin separate lines rather than a text style, so this detects those lines and correlates them with the nearby text to mark it as underlined. Detection granularity depends on how the source PDF groups its text — enable "Preserve underlined text" in the options to turn it on.
Can I convert only some pages?
Yes. Enter a page range like "1-3, 7, 10-12" to convert only those pages, or leave it blank to convert the whole document.
What happens with a password-protected PDF?
You'll be prompted to enter the password. Once verified, conversion proceeds normally — it no longer just fails with an error.
Are tables and images preserved?
Tables and images can be included when the relevant options are enabled. The converter attempts to detect aligned text columns as Word tables and extract supported images into the DOCX file. Results may vary depending on the PDF structure, image format, and page layout, so review the converted document before final use.
How are my PDF files processed?
PDF text extraction, OCR, and Word document generation run in your browser. The converter is designed to process the selected PDF locally rather than sending the document to a conversion server. You should still review your browser network activity and site integrations if your document contains highly sensitive information.