Client-Side Processing · Auto Page Sizing (A4 / Letter / Legal)

PDF to Word Converter — editable .docx with OCR and dynamic page geometry in your browser

Converts digital and scanned PDFs, preserving page dimensions (A4, Letter, Landscape), bold, italic, headings, headers, and footers — before your file downloads.

Page dimensions, tables, and images are preserved. Word page sizes and orientations match your PDF automatically. Review your converted DOCX in Microsoft Word or LibreOffice for final document sign-off.

See how it works
Dynamic A4/Letter/Landscape Sizing Bold, italic & real Heading styles Password Unlock Included Zero Server Uploads
Built & maintained by Bhavin J. Sheth | Updated

The conversion console

Upload a PDF, customize document setup, then convert.

Browser-based processing

PDF parsing, page dimension extraction, OCR, and Word document generation run securely inside your browser.

PDF to Word

Drop your PDF here

or browse files

Preparing...

PDF page

Extracted text

Extracted text preview


                

The exported DOCX file retains matched page sizes, margins, font styles, and header/footer configurations.

What this converter supports

Text extraction and formatting capabilities

Designed to extract and structure PDF text reliably while being transparent about current document limitations.

What it supports

  • Runs OCR (Tesseract) on scanned pages, with automatic detection of which pages need it.
  • Detects bold, italic, and heading-sized text from the PDF's own fonts and carries them into real Word formatting.
  • Builds a navigable Word outline from detected headings.
  • Prompts for a password on protected PDFs instead of failing.
  • Lets you convert a page range, add a header/footer/title page, and choose font, spacing, and margins.
  • Detects tables, underlined text, and images and turns them into editable Word tables, underline formatting, and embedded pictures.

Current limitations

  • Embedded images are placed in reading order, not at pixel-exact positions due to flowing layout differences.
  • Underline detection granularity depends on how text runs are grouped in the source PDF.
  • Paragraphs you edit in the correction preview lose bold/italic/heading formatting as plain text is substituted. Editing OCR text may simplify some original formatting.
  • Only one PDF at a time is processed in the current single-file flow.

From PDF to editable Word document

Upload, configure, then convert — password-protected files just add one extra step.

1

Upload your PDF

Drag it onto the console, or click Select PDF File. If it's password protected, you'll be asked to unlock it first.

Scanned documents work too — OCR runs automatically where needed.
2

Set your options

Configure exactly what you need before converting:

Pages & layout: a page range, title page, font, spacing, margins.
Header/footer & OCR: page numbers, custom text, OCR language.
3

Convert

Click Convert to Word. Enable side-by-side view first if you want to watch each page's extracted text as it's processed.

4

Review and download

Check the preview, then choose a file name and format — .docx, .md, or .txt.

Real Heading styles Browser processing

What this converter is suited for

And where document complexity might require additional review.

Situations where the PDF to Word converter fits, and situations needing more caution
Use caseWhyFit
Editing a scanned contractOCR turns scanned pages into editable text, with bold/italic and headings preserved.Strong fit
Repurposing a report's textHeadings become a real, navigable Word outline; body text keeps its formatting.Strong fit
Converting just a few pagesPage range selection avoids converting an entire long document.Strong fit
A document full of data tablesEnable "Detect tables" to convert detected aligned columns into editable Word tables.Good fit
A brochure with photos or diagramsEnable "Extract and embed images" to bring pictures into the output, placed in reading order.Good fit
Converting dozens of PDFs at onceOnly one file is processed at a time currently.Not yet a fit
Built features and roadmap

What's included in this version

Everything marked "supported" is built and functional. Everything marked "planned" is kept separate for future releases.

Password-protected PDFs Supported

You're prompted for the password and conversion proceeds normally once verified, instead of failing outright.

Page range selection Supported

Enter a range like "1-3, 7, 10-12" to convert only those specific pages instead of the entire document.

Header on every page Supported

Custom text, file name, or today's date placed at the top of every page via OOXML headers.

Font, spacing & margin options Supported

Pick a font family, base size, and compact/normal/wide presets for line spacing and document margins.

Title page Supported

An optional cover page with document title and date ahead of the main content.

Italic preservation Supported

Detected from font information and carried into the Word file as real italic formatting.

Heading detection Supported

Distinct larger font sizes are mapped to Word Heading 1/2/3 styles for a navigable outline.

Side-by-side view Supported

Watch rendered PDF pages next to extracted text as conversion progresses page by page.

Dark mode for the console Supported

A toggle for the converter interface, persisting your preference locally.

.txt and .md export Supported

Export as Markdown syntax or plain text alongside standard .docx files.

Batch conversion + ZIP Planned

Requires a dedicated multi-file queue with per-file progress tracking in future updates.

Table detection Supported

Aligned columns of text are parsed and converted into Word tables based on spatial positioning.

Image extraction & embedding Supported

Images are extracted and embedded in Word files in reading order per page.

OCR image preprocessing Supported

Optional noise reduction and contrast binarization run before OCR processing.

Underline preservation Supported

Detects thin drawn lines in PDFs and correlates them with nearby text runs.

Editable OCR-correction preview Supported

Review and edit extracted text prior to final document generation.

Troubleshooting common conversion issues

Understand why issues happen and how to resolve them quickly.

Symptom
Why it happens
Fix

"No text could be extracted"

The PDF is scanned and OCR is set to Off, or the OCR language doesn't match the document's actual language.

Switch OCR mode to Auto or Always, and check the language setting.

A table looks like scrambled text

The "Detect tables" option wasn't enabled for this conversion, so columns were extracted as plain wrapped text.

Rebuild the table manually, or re-convert with "Detect tables" turned on.

Some text isn't bold/italic anymore

Detection relies on font metadata in the PDF. Scans or flattened files often lack explicit style names.

Reapply formatting manually in Word where needed; this is a source-file limitation.

"Incorrect password" on a password I'm sure is right

PDF passwords are case-sensitive. Some PDFs also have a separate permissions password, which is different from the password required to open the file.

Double check capitalization and ensure you are entering the user/open password.

OCR is slow

OCR runs entirely in your browser and takes processing time per page, especially during initial language downloads.

This is expected behavior — set OCR mode to Auto so it only runs on scanned pages.

"No text could be extracted"

Why: OCR off, or scanned PDF.

Fix: Set OCR to Auto/Always, check language.

Table looks scrambled

Why: "Detect tables" wasn't enabled.

Fix: Re-convert with it turned on.

Bold/italic missing

Why: Source PDF lacks font metadata.

Fix: Reapply manually in Word.

"Incorrect password"

Why: Case-sensitive or permissions password used.

Fix: Confirm exact open password.

OCR is slow

Why: Runs locally in browser.

Fix: Normal — use Auto mode.

Best practices for clean conversions

Recommended approaches for achieving optimal conversion results.

Set the OCR language correctly

Tesseract accuracy relies on matching the language setting to the document's actual language.

Use a page range for long documents

Converting specific pages is faster and easier to review when you only need a section of a large PDF.

Turn on side-by-side view for scans

Comparing extracted text directly against the rendered page helps spot OCR discrepancies instantly.

Check heading outlines in Word

Open Word's Navigation pane after converting to verify that headings align with document structure.

Keep original PDFs as backup

Retain source files until complex tables or specialized layouts are fully verified in the DOCX output.

Export as Markdown for wikis

When preparing documentation for Markdown-based platforms, the .md export retains formatting directly.

PDF to Word FAQs

Find answers about scanned PDFs, formatting, page ranges, passwords, privacy, and converted Word documents.

Can I convert scanned PDFs to editable Word documents?

Yes. This converter uses Tesseract.js OCR in your browser to recognize text from scanned pages. In Auto mode, pages with little or no extractable text are sent through OCR. Results depend on scan quality, document layout, and the selected language, so review the converted document before using it.

Does this preserve formatting like bold, italic, and headings?

Bold and italic are detected from the PDF's own font information and carried into the Word file as real character formatting. Larger, distinct font sizes are detected as headings and mapped to Word's built-in Heading 1/2/3 styles, which gives the document a real, navigable outline in Word — not just bigger text.

Does this preserve underlined text?

Yes. Underlines in PDFs are usually drawn as thin separate lines rather than a text style, so this detects those lines and correlates them with the nearby text to mark it as underlined. Detection granularity depends on how the source PDF groups its text — enable "Preserve underlined text" in the options to turn it on.

Can I convert only some pages?

Yes. Enter a page range like "1-3, 7, 10-12" to convert only those pages, or leave it blank to convert the whole document.

What happens with a password-protected PDF?

You'll be prompted to enter the password. Once verified, conversion proceeds normally — it no longer just fails with an error.

Are tables and images preserved?

Tables and images can be included when the relevant options are enabled. The converter attempts to detect aligned text columns as Word tables and extract supported images into the DOCX file. Results may vary depending on the PDF structure, image format, and page layout, so review the converted document before final use.

How are my PDF files processed?

PDF text extraction, OCR, and Word document generation run in your browser. The converter is designed to process the selected PDF locally rather than sending the document to a conversion server. You should still review your browser network activity and site integrations if your document contains highly sensitive information.