← All guides

Document structure

What's Actually Inside a Word Document?

Before you can understand why a Word to PDF converter sometimes gets a document wrong, it helps to know what a .docx file actually is under the hood.

Open a .docx file in a text editor and you won't see readable text — you'll see binary. That's because a modern Word document isn't really a single file at all. It's a ZIP archive containing a small filesystem of XML files, a format Microsoft calls Office Open XML (OOXML). Rename any .docx to .zip and you can unzip it to see this for yourself.

The pieces inside the archive

Inside that archive, a few files do most of the work. document.xml holds the actual text and paragraph structure. styles.xml defines every heading style, font, and paragraph format the document uses. A media folder holds embedded images. A web of relationships files ties all of these pieces together — which image goes where, which style applies to which paragraph, where a hyperlink points.

This matters because a Word document isn't really "text with formatting bolted on." The structure and the content are woven together. A heading isn't just big, bold text — it's a paragraph tagged with a "Heading 1" style, which is itself defined elsewhere in the file and might be referenced by a table of contents field somewhere else again.

The structural elements that make a document a document

  • Styles and formatting — fonts, sizes, colors, and spacing are usually inherited from named styles, not applied character by character.
  • Sections and page setup — margins, orientation, and page size can change partway through a document via section breaks, which is how one document can mix portrait and landscape pages.
  • Headers, footers, and page numbering — often different on the first page, odd pages, and even pages, all tracked separately.
  • Tables — with their own row heights, merged cells, and nested formatting rules that are notoriously easy for a converter to mishandle.
  • Embedded objects — images, charts, and OLE objects (like an embedded spreadsheet) that need to be rendered, not just copied.
  • Fields — dynamic content like a table of contents, page count, or last-modified date, computed at render time rather than stored as plain text.

Why this matters when you convert to PDF

A PDF, unlike a .docx, is a fixed-layout format — every piece of text has an exact coordinate on an exact page. So somewhere in the conversion process, all of that flowing, style-dependent, field-computed structure has to be resolved into fixed positions. That resolution step is where cheap or naive converters fall apart: a merged table cell renders wrong, a section break gets ignored and page orientation flips unexpectedly, or a font that isn't installed on the converting server gets silently substituted for something that changes the whole layout.

A reliable Word to PDF converter has to actually understand this structure — parsing the styles, resolving the fields, laying out the tables — rather than guessing at what the document should look like from surface-level text. That's the difference between a PDF that looks like your document and one that's merely a rough approximation of it.

See it handled properly — drop in a real .docx and check the layout yourself.

Try the converter

More guides

Conversion

What happens when you convert Word to PDF

Comparison

PDF vs Word: why PDF wins for sharing

Security

Document security 101

Free conversion

How to convert Word to PDF for free

Troubleshooting

Common Word to PDF problems, fixed

File size

Why your PDF is huge (and how to shrink it)