Word (DOCX) to PDF Converter
Turns a .docx into a PDF in your browser, keeping headings, paragraphs and lists, and saying plainly what it left behind.
This tool needs JavaScript: the document is opened and converted in your browser, and it is never uploaded.
72 points is an inch. 56 is about 2cm, which is the usual document margin.
Reading the document…
No document yet. Headings, paragraphs and lists come across; layout, images and footnotes do not, and the tool says which.
Turns a .docx into a PDF without uploading it. Documents are the category where uploading is least acceptable — a CV, a contract, a letter from a solicitor — and none of it needs to leave the tab.
What comes across: the text, the heading levels, the lists, the block quotes, the order. What does not: the layout. This is a conversion of the content, not a reproduction of the page, and the output says which parts it left behind rather than letting you find out later.
How to use
Choose the .docx, or drop it on the box. Pick a page size and margin, then download.
Example
report.docx
34 blocks · 3 pages · 12.4 KB
1 table flattened into lines, 2 images left out.
The preview under that shows the blocks in order with their heading levels, so you can see what the PDF will contain before downloading it.
Pitfalls
This is not Word’s own PDF export. Word reproduces the page: fonts, columns, floats, headers, footers, page numbers. This reproduces the text and its structure. For a document where the layout is the point, open it in Word or LibreOffice and print to PDF.
Latin script only. The PDF names the fonts every reader already has, which cost nothing and work everywhere, and those cover Latin script with WinAnsi encoding. Greek, Cyrillic, Hebrew, Arabic and CJK cannot be written with them: those characters come out as question marks, and the tool counts them and says so. Embedding a font would fix it and would be a different, much larger tool.
Images are left out, not lost. They are still in your .docx. The count is reported so you know how many to add back.
Tables are flattened into lines, cells joined with a separator. A two-column table usually reads acceptably that way; a spreadsheet pasted into Word does not, and that is worth knowing before you send it.
Headers, footers, footnotes and page numbers are not read. They live in separate parts of the archive and belong to the page layout rather than the content.
Tracked changes come through as accepted text, because the deletions are marked up separately and this reads the current text. Accept or reject them in Word first if that matters.
A .doc is not a .docx. The pre-2007 format is a binary nobody should be parsing in a browser. Open it and save as .docx first; the tool says so rather than failing silently.
Compatibility
Everything happens in your browser. The .docx is opened with a zip reader written for this, the text is
read out of word/document.xml, and the PDF is written by hand. Nothing is uploaded and nothing is
stored.
A .docx is a zip of XML: paragraphs are w:p, runs are w:r, the style is in w:pStyle and lists are
marked by w:numPr. Those are the parts that are stable and documented, and they are the parts this
reads. Word’s own style names are mapped to heading levels, lists and quotes, so a document written with
the built-in styles keeps its structure and one formatted by hand with bold text does not.
The PDF names the base fourteen fonts — Helvetica and Courier, in bold and oblique — with WinAnsiEncoding. Line breaking is measured with your browser’s own Helvetica, which is metrically compatible with the one the reader will use, so the wrapping in the file matches what was measured.
Limits: 25MB for the file. The cross-reference table is checked by the test suite, every offset followed to its object, because a wrong one produces a PDF that opens in some readers and not others.