Convert a PDF and the emoji turn into boxes, the quotes become question marks, and the dashes break. Here's what happens to special characters and how to rescue them.
You convert a report that has a few emoji, some smart quotes, and an en dash in a date range. The Word file comes back with little boxes where the emoji were, curly quotes turned into question marks, and the "2018–2021" range now reads "2018?2021." You've just hit the special-character wall — the single most common reason a converted document needs manual fixing. The good news: it's predictable, and most of it is fixable before you start editing. Here's what actually happens under the hood.
PDFs store glyphs, not text — each character is drawn as a shape, and the mapping back to Unicode is only as good as the original file's embedded font information. Emoji are the worst case: they're multibyte Unicode, frequently missing from the font the PDF used, and the converter has to decide what to do with a glyph it can't identify. The result is the classic replacement box, or a drop to a fallback font that changes the whole line's spacing. The PDF to Word converter handles the standard set — Latin letters, common punctuation — reliably; the trouble starts where the document got creative with its characters.
Three families cause ninety percent of the damage. Smart quotes and apostrophes (U+2018–U+2019) break when the source font maps them oddly and the converter falls back to ASCII — they arrive as straight quotes or question marks. Dashes — the en dash and em dash — get mistaken for hyphens or split into "?" because they're separate glyphs from the hyphen key. And any character outside the document's declared encoding, from a bullet to a currency symbol, is a lottery. The fix for all three is the same: after converting, run a find-and-replace pass for the specific characters you know were in the original. Pass the whole thing through the text polish tool afterwards and it will catch the strays you missed — the doubled spaces, the orphaned punctuation, the artifacts sitting mid-sentence.
If the PDF is a scan rather than a born-digital file, the character errors aren't a font issue — they're an OCR misread. An OCR engine reads "2018–2021" and occasionally emits "2018-2021" or worse, especially in small print. The counter-intuitive part: for scanned documents, don't try to fix every misread character by hand. Instead, verify the fields that matter — numbers, dates, email addresses, prices — and use the rest of the document as a working draft. If the document needs to be clean enough to publish, retype the critical lines rather than trusting the OCR, and run the article generator only on the structure, not the recovered text. We covered when a conversion is even worth doing in our guide to when you should NOT convert a PDF — a heavily typographic brochure is a good candidate for staying a PDF. For everything else, convert, sweep for the three character families, and polish the residue away.
PDF to Word
Convert PDF to editable Word (.docx) free — no watermarks, no registration. Smart text extraction preserves headings, paragraphs, and formatting. Auto-detects and converts PDF tables. Scanned PDF support with Google Cloud Vision OCR text extraction. Embedded images preserved in output.
Text Polish & Rewrite
Polish, rewrite, shorten, or expand your text with AI.
AI Article Generator
Generate complete, well-structured articles from a topic and keywords with AI.