PDF to Word Converter
Pulls the text out of a PDF in your browser and writes it into an editable .docx, and says plainly what a PDF cannot give back.
This tool needs JavaScript: the PDF is opened and read in your browser, and it is never uploaded.
Join rejoins the lines into paragraphs, ending one at a full stop. Lines keeps the breaks the PDF had, which suits verse, addresses and code.
Reading the PDF…
No PDF yet. The text comes back; the layout does not, because a PDF does not store one. What you get is the words, in order, as editable paragraphs.
Converting out of PDF is a recovery job, not a conversion, and every tool that pretends otherwise is guessing. This one gets the text back and says what it left behind.
A PDF does not store a document. It stores instructions for painting one: pick a font, move to a position, paint these characters. There are no paragraphs in the file, no headings, no reading order beyond the order the generator happened to emit. So the text can be recovered exactly, and the structure cannot be recovered at all. What comes out here is the words, grouped into lines by the vertical moves between them and into paragraphs by the punctuation, in a .docx you can edit.
The file never leaves your device. Every “PDF to Word” site works by uploading the file to a server, and the documents people convert are payslips, contracts, medical letters and bank statements. This one opens the PDF in the tab and writes the Word file in the tab.
How to use
- Choose a PDF, or drop one on the box.
- Read the per-page character counts before you download anything: a page showing “no text” is a picture of a page, and nothing will come out of it.
- Pick whether paragraphs are rejoined or the original line breaks are kept, then download the .docx.
Example
A four-page contract, 11,400 characters:
4 pages · 11,412 characters · 96 paragraphs · 31.4 KB
Page 1: 3,104 characters
Page 2: 3,380 characters
Page 3: 2,918 characters
Page 4: 2,010 characters
What comes across the text, as editable paragraphs
What does not fonts, sizes, columns, tables, images, headers and footers
Why a PDF stores those as positions on a page rather than as structure
And the same tool on a scanned page:
1 page · 0 characters
No text came back. The pages this could not decode are images, which is what a
scanned document is: the characters are not in the file, only a picture of them.
That second result is the honest one. A converter that produced a Word file full of nothing there would have wasted your time twice.
Pitfalls
A scan is a photograph. If the PDF came from a scanner, a phone camera or a fax, the text is not in the file. Getting it out needs optical character recognition, which is a different job and a much larger piece of software. The per-page counts here tell you which kind you have before you download anything.
Layout does not survive, and no tool makes it survive. Columns, tables, text boxes and pull quotes are positions in a PDF, not structures. Reading them back in the right order is guesswork, and the guesses are wrong often enough that the honest output is a single column of text.
Headings are not marked. A heading in a PDF is larger type at a position. This deliberately does not guess which lines were headings, because a document that looks structured and is structured wrongly is harder to fix than plain paragraphs.
Encrypted PDFs are refused. A PDF with any password, including the empty owner password that only restricts printing, has its streams scrambled. Removing that protection is not something this does.
Ligatures and odd fonts can come back wrong. A font that carries no ToUnicode mapping has no way to say which characters its codes mean. Most generators include one; a few subset fonts do not, and their text comes out as nonsense rather than silently as something plausible.
Hyphens at line ends stay hyphens. Rejoining “docu-” and “mentation” would sometimes be right and sometimes destroy a real hyphen, so the line break is removed and the hyphen is left for you to see.
The .docx is text, so it is small. A 4MB PDF becomes a 30KB Word file. That is not a compression trick: the images and the fonts are what made the PDF large, and neither is in the output.
Compatibility
Everything runs in the browser: the PDF is read with a parser written for this, and the .docx is written as the zip of XML that the format actually is. Nothing is uploaded and nothing is stored.
Flate is the compression PDFs use for text, and the browser’s own DecompressionStream handles it, so
there is no library to download. That needs Chrome 80, Edge 80, Safari 16.4 or Firefox 113. Streams
compressed another way, and streams that are images, are counted and reported rather than skipped
silently.
The output opens in Word 2007 and later, Pages, LibreOffice, Google Docs and anything else that reads .docx. Headings carry outline levels, so Word’s navigation pane and a generated table of contents both work if you apply heading styles yourself afterwards.
Files up to 30MB, which is a long report. The reading is a scan of the file rather than a render of the pages, so it takes about as long as reading the file off disk.