PDF to Word Converter

Pulls the text out of a PDF in your browser and writes it into an editable .docx, and says plainly what a PDF cannot give back.

This tool needs JavaScript: the PDF is opened and read in your browser, and it is never uploaded.

Line breaks

Join rejoins the lines into paragraphs, ending one at a full stop. Lines keeps the breaks the PDF had, which suits verse, addresses and code.

No PDF yet. The text comes back; the layout does not, because a PDF does not store one. What you get is the words, in order, as editable paragraphs.

Converting out of PDF is a recovery job, not a conversion, and every tool that pretends otherwise is guessing. This one gets the text back and says what it left behind.

A PDF does not store a document. It stores instructions for painting one: pick a font, move to a position, paint these characters. There are no paragraphs in the file, no headings, no reading order beyond the order the generator happened to emit. So the text can be recovered exactly, and the structure cannot be recovered at all. What comes out here is the words, grouped into lines by the vertical moves between them and into paragraphs by the punctuation, in a .docx you can edit.

The file never leaves your device. Every “PDF to Word” site works by uploading the file to a server, and the documents people convert are payslips, contracts, medical letters and bank statements. This one opens the PDF in the tab and writes the Word file in the tab.

How to use

  1. Choose a PDF, or drop one on the box.
  2. Read the per-page character counts before you download anything: a page showing “no text” is a picture of a page, and nothing will come out of it.
  3. Pick whether paragraphs are rejoined or the original line breaks are kept, then download the .docx.

Example

A four-page contract, 11,400 characters:

4 pages · 11,412 characters · 96 paragraphs · 31.4 KB

Page 1: 3,104 characters
Page 2: 3,380 characters
Page 3: 2,918 characters
Page 4: 2,010 characters

What comes across   the text, as editable paragraphs
What does not       fonts, sizes, columns, tables, images, headers and footers
Why                 a PDF stores those as positions on a page rather than as structure

And the same tool on a scanned page:

1 page · 0 characters

No text came back. The pages this could not decode are images, which is what a
scanned document is: the characters are not in the file, only a picture of them.

That second result is the honest one. A converter that produced a Word file full of nothing there would have wasted your time twice.

Pitfalls

A scan is a photograph. If the PDF came from a scanner, a phone camera or a fax, the text is not in the file. Getting it out needs optical character recognition, which is a different job and a much larger piece of software. The per-page counts here tell you which kind you have before you download anything.

Layout does not survive, and no tool makes it survive. Columns, tables, text boxes and pull quotes are positions in a PDF, not structures. Reading them back in the right order is guesswork, and the guesses are wrong often enough that the honest output is a single column of text.

Headings are not marked. A heading in a PDF is larger type at a position. This deliberately does not guess which lines were headings, because a document that looks structured and is structured wrongly is harder to fix than plain paragraphs.

Encrypted PDFs are refused. A PDF with any password, including the empty owner password that only restricts printing, has its streams scrambled. Removing that protection is not something this does.

Ligatures and odd fonts can come back wrong. A font that carries no ToUnicode mapping has no way to say which characters its codes mean. Most generators include one; a few subset fonts do not, and their text comes out as nonsense rather than silently as something plausible.

Hyphens at line ends stay hyphens. Rejoining “docu-” and “mentation” would sometimes be right and sometimes destroy a real hyphen, so the line break is removed and the hyphen is left for you to see.

The .docx is text, so it is small. A 4MB PDF becomes a 30KB Word file. That is not a compression trick: the images and the fonts are what made the PDF large, and neither is in the output.

Compatibility

Everything runs in the browser: the PDF is read with a parser written for this, and the .docx is written as the zip of XML that the format actually is. Nothing is uploaded and nothing is stored.

Flate is the compression PDFs use for text, and the browser’s own DecompressionStream handles it, so there is no library to download. That needs Chrome 80, Edge 80, Safari 16.4 or Firefox 113. Streams compressed another way, and streams that are images, are counted and reported rather than skipped silently.

The output opens in Word 2007 and later, Pages, LibreOffice, Google Docs and anything else that reads .docx. Headings carry outline levels, so Word’s navigation pane and a generated table of contents both work if you apply heading styles yourself afterwards.

Files up to 30MB, which is a long report. The reading is a scan of the file rather than a render of the pages, so it takes about as long as reading the file off disk.

Frequently asked questions

Why is my scanned document coming out empty?
Because there is no text in it. A scan is an image of a page, and the characters exist only as pixels. That needs OCR, which reads the picture and guesses the letters. This tool reads the file and reports exactly what is in it, which is why it tells you the page has no text rather than producing an empty document.
Can I get the original layout back?
No, and neither can the paid tools, whatever their marketing says. They guess at columns and tables from the positions of the characters, and the guesses fail on anything unusual. Text in one column is the result that is reliably correct.
Is my file uploaded anywhere?
No. The conversion happens in this tab, and if you disconnect from the network after the page has loaded it still works. That is worth caring about for the documents people usually convert.
What about tables?
A table in a PDF is a set of text positions and, sometimes, some lines drawn behind them. The cells come back as text in reading order rather than as a table. Rebuilding the grid means guessing the column boundaries from the x positions, and that guess is wrong often enough to be worse than plain text.
Why does the text have the line breaks from the PDF?
Only if you ask for it. By default, lines are joined into paragraphs and a paragraph ends at a full stop, which is what you want for editing. Keeping the original breaks is useful for verse, addresses and code, so it is an option.
Weekly drops

New tools, when there are new tools

One email when something worth using ships. No schedule to fill, so no filler.

Your address goes nowhere else, and one click unsubscribes.