Guide
What an applicant tracking system actually reads from your resume
Written from the inside of a resume parser: what a PDF gives up and what it silently loses, why bold text is invisible to it, how two columns interleave, and the five things that reliably break extraction.
Last updated
How do we know this?
Because we wrote a parser. Haisleaf imports resumes from PDF and DOCX, which means extracting text from both formats and reassembling it into sections and entries. Everything on this page is something that parser can or cannot see. It is not a list of tips; it is what the inside of the problem looks like.
Every applicant tracking system is different, and none of them publishes its parser. But they all start from the same place — pulling text out of a file — and the things that break at that stage break everywhere.
What survives, and what is lost
The two formats are not two grades of the same thing. They know genuinely different facts, and neither is a degraded version of the other.
| From a PDF | From a DOCX | |
|---|---|---|
| The words themselves | Yes | Yes |
| Reading order | Usually — see columns below | Yes |
| Which lines are headings | No. Inferred from size and position | Yes, from the style |
| Bold and italic | No | Yes |
| Bullet-list structure | Inferred from the marker glyph | Yes |
| Where each line sits on the page | Yes, to the pixel | No |
| Anything inside an image | No | No |
The row that surprises people is bold and italic. A PDF does not hand a parser the name of the font a run of text is set in — the library everyone uses returns an internal identifier and a generic family like “serif”, even when the document’s actual font is Arial-BoldMT. So a parser reading your PDF cannot tell that your job titles are bold. If emphasis is the only thing marking something as a heading, that information does not exist by the time anything reads it.
Why do two columns come out scrambled?
Because a PDF has no idea it has columns. It is a list of text fragments, each with a position on the page. Reading order has to be reconstructed from those positions, and the obvious reconstruction — top to bottom, left to right, row by row — reads straight across the gutter.
So a skills sidebar beside your experience can interleave: one line of a job, one skill, one line of a job. The text is all there and it is in an order nobody can use. Deciding which x-positions are a column and which are a wide space is a judgement about the whole page, and a parser that gets it wrong gets it wrong silently.
What to do: if a system is parsing rather than a person reading, a single column is the safe answer. Two columns are not fatal — plenty of parsers handle them — but they are the difference between a format that always works and one that usually does.
The five things that reliably break extraction
In rough order of how badly they break it, and all five are things we have watched go wrong on real files:
- A resume that is really a picture. A design exported as an image, or a scan, contains no characters at all. A parser gets an empty document — not a bad parse, nothing. If you can’t select the text with a cursor, neither can it.
- Tables used for layout. Cell order and visual order are not the same thing, so the text arrives shuffled in a way that looks arbitrary.
- Letter-spaced headings. A heading tracked out as EXPERIENCE comes through with a space between every letter — E X P E R I E N C E — because the gaps really are wider than the space character. It stops matching the word.
- Dates only a human can read. “Summer 2019” and “Ongoing” are not dates any parser will turn into a range. A month and a year, or a year, parse everywhere.
- Contact details in a header or footer. Page furniture is often extracted separately from body text, or last, or not at all. Your email belongs in the document.
Does a PDF or a Word file work better?
A DOCX carries more structure — it says outright which lines are headings and which are list items, where a PDF makes a parser guess from size and position. If a posting asks for Word, send Word.
Otherwise send a PDF, and send one whose text is selectable. A PDF is the only format that looks the same everywhere, and a Word file re-flowed by someone else’s copy of Word is a document you have not seen. The structure a DOCX would have carried is worth less than knowing what arrived.
How can I see what a parser sees?
Open your PDF, select all, and paste it into a plain-text editor. That is approximately the input. Read it in that order and ask whether it still makes sense — if the columns interleaved or a heading lost its letters, you will see it immediately.
Haisleaf does this for you: every resume exports as plain text as well as PDF, from the same document, so you can check the machine-readable version of the exact file you are about to send. It is free to try — the plain-text export is on every plan.
What about keywords?
Keyword matching is a separate step and a later one: it happens after extraction, on the text that came out. It is worth doing — a posting’s own vocabulary is the vocabulary it will be screened against — but it is worth nothing if the extraction lost half your resume first. Get the text out cleanly, then worry about the words in it.
One thing worth saying plainly, because plenty of advice says otherwise: there is no score any tool can give you that an employer will see. A match score is a way of noticing a word you have not used. It is advice about a document, not a verdict on a person.
The other half of the format question is length — why your resume runs onto a second page covers what actually decides where the break lands.