How to open a .docx without Word, and get the text out when Word says it is corrupt

Two situations: you have a .docx and no Word, or Word refuses the file with "Word found unreadable content". Below are the free programs that open it, what each asks of you, and how to get the text out of a file none of them will open.

Ashton
File formats

Either way, what you usually need right now is the words, not the layout.

No Word on this computer: free ways to open a .docx

Maybe the Microsoft 365 subscription lapsed, Word is only on the work PC, you are on a Mac or a phone, or your school account ended. First check one thing. Microsoft's licensing documentation says unlicensed Microsoft 365 Apps stay installed in "reduced functionality mode", where "users can only view and print their documents." So if Word is still there, double-clicking the file may still open it for reading.

Or the files move to a new computer that never had Word. A Mac user on Apple Community, after retiring:

When I retired, I put personal files on a flash drive. I can download the files, but the doc and docx files will not open. "Cannot open files at this time". I tried opening the files with the LibreOffice, did not work. …

April 2021, MacBook Pro on macOS 10.15.

Apple Community

LibreOffice, the usual first suggestion, did not help here. Another reply pointed to Pages: "Pages can open word documents." Here is the whole list:

OptionWhat it asks of youWorth knowing
Word for the webA free Microsoft account and a browserFree on the web from Microsoft. Open your file with "Upload a file". Free accounts get 5 GB of OneDrive.
Google DocsA Google accountUpload to Drive (New > File upload), then double-click. Takes .docx and .doc newer than Office 95, up to 50 MB when converted.
LibreOffice 26.8An install on Windows, macOS or LinuxFree, open source, released 26 August 2026. Reads and writes DOCX, offline.
PagesA Mac, iPad or iPhoneFile > Open. Opens .doc and .docx. Apple says Pages stays free; the paid Creator Studio adds templates, media and AI features.
TextEditNothing, it comes with macOSApple says it opens documents from other word processors, including Microsoft Word.
DOCX to text on this siteNothingText, Markdown or HTML only. See below.

On Windows you will still find old advice to open the file in WordPad. That does not work on an up-to-date PC: Microsoft removed WordPad from all editions starting with Windows 11 version 24H2.

For work files: Word for the web and Google Docs put the document into a personal cloud account, which your employer or client may not allow. LibreOffice, Pages and TextEdit keep it on your computer.

"Word found unreadable content": what the messages mean

Word usually refuses a file in a chain of messages. A Microsoft Q&A thread that 300+ people marked as the same question records the sequence:

What Word saysWhat it tells you
The file cannot be opened because there are problems with the contents.Word could not read part of the file. Click Details before you close it.
Word found unreadable content in "file name". Do you want to recover the contents of this document? If you trust the source of this document, click Yes.Word is offering its own repair. Clicking Yes is the only way to find out whether it works.
Microsoft Office cannot open this file because some parts are missing or invalid.The repair did not get past the damage.

Click Details, too. In another thread it showed "The name in the end tag element must match the element type in the start tag Location: Part:/word/document.xml, Line: 2, Column: 3479". The part holding the text has a broken tag at that exact spot, which matters in the next section.

If you have Word, Microsoft's article on damaged documents suggests two more things. File > Open, click the file once, click the arrow on the Open button, choose Open and Repair. Or set the file type box in that dialog to Recover Text from Any File(*.*): formatting, graphics and fields are lost, while field text, headers, footers, footnotes and endnotes are kept as plain text.

Another program sometimes gets further, because it reads the file its own way. A reply on the same thread, from 2010:

Recently, in a class of 20 students, it has occurred 5 times for students when creating documents with a Table of Contents or Index. None of the fixes at the support link work. Only opening the file with Open Office allows recovery of some items.

A reply in "File cannot be opened because there are problems with contents", 300+ people with the same question.

Microsoft Q&A

Today that would be LibreOffice, which grew out of OpenOffice. Expect "some items", not the whole file.

A .docx is a zip file: getting the text from word/document.xml

Microsoft's Open XML documentation says it too: rename a .docx to .zip to see its parts, and document.xml "contains the content of the main body of the document". On the broken end tag thread (50+ people with the same question), a volunteer answered:

A Word 2007 docx file is nothing more than a zip package. If you change the extension from .docx to .zip, any zip archiver will be able to open it.

Answer to "The office Open XML file cannot be opened because there are problems with the content", July 2010.

Microsoft Q&A

On a copy, never your only one:

Copy the .docx
Rename the copy to .zip
Extract it
Open word/document.xml in a text editor
Read what sits between <w:t> and </w:t>
If you cannot see ".docx" at the end of the name, extensions are hidden; turn them on in File Explorer.

Text sits inside <w:t> elements, so a long document is slow going, but nothing is hidden. In our test files the pictures were in word/media as ordinary image files, and headers, footers and footnotes had their own XML files.

Two limits. If extracting fails, the zip is incomplete (often an attachment or download that never finished) and the text went with it; get the file again and compare sizes with the sender. And a broken tag at a known line and column can be fixed by someone comfortable with XML, who then zips the parts back up. That is the real repair, and not a job for everyone.

Only need the text? Extract it in your browser

The DOCX to text tool on this site is for the "I only need the words" case. It reads one .docx at a time inside your browser, so nothing is uploaded, and it works in a phone browser (big files are slower). Pick plain text, Markdown or HTML, check the preview, then copy or download. What each output keeps, from reading its code and running test files through it:

In the documentPlain textMarkdownHTML
Paragraphs and headingsKept, formatting droppedKept, with # headings, bold, italics and linksKept as tags
TablesEach cell on its own lineA pipe tableA real table
PicturesDroppedEmbedded in the file as data, which can make it largeEmbedded the same way
Footnotes and endnotesDroppedListed at the endListed at the end
Headers, footers, commentsDroppedDroppedDropped
Tracked changesInsertions kept, deletions droppedSameSame

It is not a repair tool. We ran damaged test files through its parser (mammoth 1.12.2). A file cut off partway failed with "The file is cut off, or its format does not match the extension." A file with an unbalanced tag, like the Q&A case, failed too: the parser stops at the first XML error. A file with valid XML and an element it did not expect gave up all its text. So it helps when the file is intact and a program is fussy, not when bytes are missing.

It cannot read old .doc files or password-protected .docx files either. A .docx we saved with a password was not a zip at all, so renaming fails too. The tool now recognises both cases: it says the file is an old Word file (.doc) or a password-protected document, and asks you to open it in Word or Google Docs, remove any password and save it as .docx.

Extract the text from a .docx

When nothing works

Sometimes every route is closed. This was posted on Microsoft Q&A in October 2025:

I have a submission in the next couple of minutes & my Microsoft Word file will not open now. It seemed to be working as expected earlier.

From "URGENT: Microsoft Word found unreadable content in (File Name)."

Microsoft Q&A

A Microsoft moderator suggested this article's routes: rename to .zip and find word/document.xml, try Google Docs or LibreOffice, start Word in safe mode. The asker replied: "None of it is working. I just got off a call with Microsoft Help. Nothing works." Later that month someone added that their Master's thesis, "over 50 pages long", had done the same, and a third-party recovery tool had failed too.

When bytes are missing, no program can put them back. Look for another copy: ask the sender again, check sent mail, a USB stick or a cloud folder. A PDF from the sender opens in any browser, and if you need to edit it as Word again, the PDF to Word tool here rebuilds its paragraphs and tables into a new .docx. It needs a PDF with real text, not a scan, and the layout is reconstructed, not copied.

Turn a PDF back into a Word file

Common questions

Does any of this work for old .doc files?
Not the extractor on this site, which reads .docx only. Google Docs converts .doc files newer than Office 95, and Pages opens .doc as well.
I need to send it back as a .docx after editing.
Use LibreOffice, which reads and writes DOCX, or Pages, which can save a copy in Word format. The extractor here only gives you text.
Renaming to .zip gives an error.
The download may be incomplete, the file may be password-protected (then it is not a zip inside), or it may be an old .doc that only has a .docx name.
Can I do this on a phone?
Pages runs on iPhone and iPad. The extractor here runs in a phone browser; very large documents take a while.

Last updated September 30, 2026 · As of September 30, 2026. Checked against LibreOffice 26.8, Pages 15.4 for Mac (Apple's user guide), Windows 11 version 24H2, and mammoth 1.12.2, the parser inside the DOCX to text tool.