Guide

Does an applicant tracking system read bold text?

The usual answer is that bold is a font-weight flag any parser reads straight through. Measured against the library most parsers are built on, a PDF hands over no font name at all — and a DOCX hands over the style by name.

Last updated

How do we know this?

Because we wrote the extractor, and because we got it wrong first. Haisleaf imports resumes from PDF and DOCX, so it has to pull text out of both. The numbers below are measured against the exact library version that code pins — not inferred from how the PDF specification reads.

One caveat worth stating plainly, because most pages on this topic state the opposite of each other without one: this describes what the pdfjs-family text API returns. That is what most resume parsers are built on, and it is not a promise about every commercial system — a vendor running its own extraction could do more work than this. The claim is about the tool nearly everybody actually uses.

What a PDF actually hands a parser

A PDF page, once parsed, is a list of text runs. For each one you get this and nothing else:

The fields a PDF text-extraction API returns for one run of text, and whether each one identifies the font.
FieldWhat it holdsTells you the font?
strThe characters of one runYes — this is the text
transformA six-number text matrixYes — position and em size
widthThe advance widthYes — where the run ends
fontNameAn internal id, e.g. g_d0_f1No — names nothing
styles[id].fontFamilyserif, sans-serif or monospaceNo — a fallback, not the font
A bold or weight flagNot present—

The two rows that decide the question are the last three. fontName looks like it should name the font and does not — it is a loaded-object id, something like g_d0_f1, meaningful only inside that one parse. And the family you can look up beside it is a fallback: it is exactly one of serif, sans-serif or monospace, chosen so a viewer has something to draw with.

Which produces the case that settles it. A page whose font is literally ABCDEF+Arial-BoldMT — a subset of Arial Bold, embedded, named in the file — reports a family of sans-serif. The word “Bold” is in the document and is not in the answer.

Is the font name in the file at all?

Yes — and this is the part the short answer compresses. The real base font name is in the PDF, in the font dictionary. It is simply not on the object the text API returns. Reaching it means asking for the page’s operator list instead: a second full parse of every content stream on the page, to recover a hint.

That is the whole economics of it. Extraction is already the most expensive thing a parser does per upload, and doubling it to learn that a job title was probably bold is a bad trade — so essentially nobody makes it. The information exists. Almost nothing pays for it.

The version of this we shipped, and deleted

Our extractor used to report whether a line was bold. It did it by testing a pattern — /bold|black|heavy/ and friends — against the font family and against the font name. Both of those are the fields above. Neither is a font name.

So the check returned false for every file it had ever seen, and had done since the day it was written. It was not a rough answer or a partial one. It was a dead branch that read like a live one, and the only reason it was caught is that somebody went and measured what those two fields actually contain instead of trusting what they were named.

It is gone now, and what replaced it reports nothing rather than reporting false — because “false” claims we looked, and for a PDF there is nothing to look at. We are telling you this because it is the same mistake the ranking advice makes: a field called fontName sounds like it names the font.

So should I stop using bold?

No. Bold does not break anything — this is not a formatting hazard like a two-column layout or a heading drawn as an image. The words underneath it come through perfectly. Bold is for the human who reads your resume after it has been parsed, and that person is the one who decides.

What to do: never let bold be the only thing carrying meaning. If a line is a section heading, let the word say so — write Experience, on its own line, in the ordinary way — rather than relying on weight to mark it. If a date matters, put it in the text, not in a styled run beside it. Then the emphasis is a bonus for the reader instead of the load-bearing structure a parser cannot see.

And if a posting asks for a Word file, send one. A DOCX names its styles, so it carries the heading structure and the emphasis that a PDF drops. The trade is covered in what an applicant tracking system actually reads.

How can I check my own resume?

Open your PDF, select all, and paste it into a plain-text editor. Everything that survives the paste is roughly what a parser gets — and everything that does not, it never had. Read it in that order and ask whether it still makes sense without a single bold word in it.

Haisleaf does this for you: every resume exports as plain text as well as PDF, from the same document, so you can read the machine-readable version of the exact file you are about to send. It is on every plan.