mypdf.space

Guide

Convert a scanned PDF to editable Word

If you have run a scanned PDF through a PDF-to-Word converter and got back a document that is blank, or that contains one enormous picture of your page, nothing is broken. The file genuinely has no text in it.

A scan is a photograph. What looks like the word "Invoice" to you is, to the file, a pattern of dark pixels. A converter can only export text that exists, so the fix is to create that text first with OCR, then convert.

Step by step

  1. 1Open OCR PDF and drop your scanned file in.
  2. 2Pick the language of the document — accuracy drops sharply if this is wrong.
  3. 3Run the recognition and wait; it processes page by page on your own hardware.
  4. 4Download the result, which is your original pages with an invisible, selectable text layer added.
  5. 5Open PDF to Word, drop in the OCR’d file, and export the .docx.
  6. 6Proofread the Word document against the original — OCR mistakes cluster around numbers, punctuation and anything handwritten.

How to tell whether your PDF is scanned

Open the file in any viewer and try to select a line of text with your cursor. If a selection highlight appears over the words, the text is real and you can convert directly. If your cursor draws a rectangle across the page instead, or selects the entire page as one object, it is a scan.

A second check: use your viewer’s search function and look for a word you can plainly see. No result means no text layer.

Getting the best OCR result

Recognition quality is decided before you start, by the scan itself. Around 300 DPI is the sweet spot — below 200 DPI characters begin to merge, and above 400 DPI you gain little but a much slower run. Straight pages matter more than people expect; a few degrees of skew from a hand-fed scanner measurably increases errors, so deskew or re-scan crooked pages if you can.

Contrast helps too. A grey photocopy of a photocopy gives the engine less to work with than a clean black-on-white original. If your source is faint, scanning again in black and white mode often beats any amount of post-processing.

Set the language correctly. An English-language engine reading a German document will guess plausible English words at every accented character, and the result is harder to fix than the original.

What the Word file will look like

Expect the words, not the design. The export carries text and basic structure across; it does not attempt to be a pixel-perfect reproduction of the original layout. Multi-column pages, tables with merged cells, sidebars and footnotes are where it simplifies most, because reconstructing those from a flat page is genuinely hard and guessing wrong is worse than not guessing.

For a letter, a report or a contract, the output is normally something you can edit immediately. For a densely formatted form or a magazine page, treat it as a way to recover the text rather than the document.

Tools used here

Frequently asked questions

  • Why is my converted Word document empty?
    Because the PDF contains images of pages rather than text. Converters export the text layer, and a scan has none. Run OCR first to create that layer, then convert.
  • Does OCR change how my pages look?
    No. The original page images are kept exactly as they are and the recognised text is added as an invisible layer positioned over them. Visually the file is unchanged; it simply becomes searchable and selectable.
  • Do I need to install anything?
    No. Both the OCR engine and the Word export run inside the browser tab, so there is nothing to install and nothing to upload. The first OCR run downloads the recognition data for your chosen language, which takes a moment.
  • How accurate is browser OCR compared to a paid service?
    On clean, straight, 300 DPI printed text it is competitive. On poor scans, unusual fonts or heavy layouts, dedicated commercial OCR is meaningfully better, and it is fair to say so — the trade you are making here is some accuracy in exchange for the file never leaving your machine.

Other guides

  • PDF too large to emailGet a file under the 25 MB Gmail limit — or any limit — without uploading it anywhere.
  • Remove a PDF passwordStrip the password from a document you have the right to open, without uploading it.
  • Redact before sharingWhy black rectangles are not redaction, and how to remove content so it stays removed.