Where AI screening quietly reintroduces bias

Samuel Reiss

Co-founder & CTO

Latest update

22 min read

IN THIS ARTICLE
The Four Fields Problem
Proxies that survive annoymization
When the rubric itself is the bias
Feedback loops, and why they hide
What monitoring actually catches
A short checklist
10 min remaining
Almost every recruiting team that adopts AI screening does the same first thing: they hide the candidate's name, photo, age and university. It feels like the hard part is done. The model can no longer see who someone is, so it can no longer prefer one group over another.

Each of those carries signal about the role. Several of them also carry signal about the person. The model cannot tell the difference, because from its position there isn't one — it sees text that correlates with outcomes in the training data and weights it accordingly.

The four fields problem

Consider what remains after a standard blind-screening pass: employment dates, job titles, postcode or city, previous employers, volunteer work, languages spoken, certifications, the software someone lists, hobbies, and the way the whole thing is written.

Each of those carries signal about the role. Several of them also carry signal about the person. The model cannot tell the difference, because from its position there isn't one — it sees text that correlates with outcomes in the training data and weights it accordingly.
The test we use internally
Take a redacted résumé from your pipeline and hand it to two colleagues. Ask them to guess the candidate's rough age, gender and whether they grew up in the country they're applying in. If your colleagues can do better than chance, so can a model — and it has thousands more examples to learn the pattern from.

Proxies that survive anonymisation

These are the ones we see cause measurable score differences most often, roughly in order of how much damage they do relative to how invisible they are:
Employment gaps. A fourteen-month gap is a strong proxy for parental leave, illness or caring responsibilities. Models trained on "continuous tenure predicts performance" penalise it heavily, and the penalty is not evenly distributed.
Postcode and commute distance. In most cities, where someone lives is a reliable proxy for both class and ethnicity. If your rubric rewards proximity to the office, you have built a geographic filter that correlates with protected characteristics.
Prestige signals in prose. Candidates coached on CV writing use a specific register — action verbs, quantified outcomes, tight parallel structure. That register is taught in some schools and networks and not others. Reward the writing and you reward the coaching.
Non-linear careers. Someone who moved into tech at 34 has a shorter tech history than someone who started at 22, and models read that as less experience rather than a later start.
Volunteer and interest lines. Religious organisations, sports, cultural associations. Almost nobody scores these deliberately, and almost every free-text model reads them.

When the rubric itself is the bias

The uncomfortable finding from our own pipelines is that most measurable disparity does not originate in the model. It originates in the criteria the model was told to apply.

A rubric that requires "8+ years in a similar role" will produce different pass rates across age groups. A rubric that weights "startup experience" heavily will favour people who could afford the salary risk of an early-stage job. A rubric that treats a named certification as mandatory rather than equivalent will filter out anyone who trained in a different country. None of this requires a biased model. A perfectly calibrated model applying these criteria produces exactly the disparity the criteria encode, faster and more consistently than a human would.

This is why we spend more engineering effort on how criteria are written than on the scoring model itself. It's also why every rubric in FairHire is versioned: when a pass-through rate shifts, the first question is always which criteria changed, not whether the model drifted.

Feedback loops, and why they hide

The slowest failure is the one that compounds. If a screening system is tuned on your past hires, it learns what your organisation has historically rewarded including the parts you're trying to change. Its recommendations then shape the next cohort of hires, which becomes the next round of training data.

Two years in, the system looks highly accurate. It predicts your hiring decisions with impressive fidelity. That accuracy is the problem: it is measuring agreement with a process, not quality of outcome.
Stage-level pass-through monitoring. A disparity that appears at one stage and disappears at the next usually points at a criterion, not a model.

What monitoring actually catches

The two criteria they changed: continuous employment history became "relevant experience, in any arrangement", and a specific cloud certification became "that certification or demonstrable equivalent". The screening stage moved from 0.71 to 0.94, and seventeen more candidates reached interview — from the same applicant pool.

The point isn't that rubric edits fix everything. It's that the disparity was visible at a specific stage, in week two, while the roles were still open. Annual reporting would have surfaced it eleven months later, attached to hires that had already been made.

A short checklist

  1. Audit the rubric before the model. Read every criterion and ask who it excludes that the job doesn't require.
  2. Test for guessability. If a human can infer demographics from a redacted profile, treat the profile as unredacted.
  3. Monitor per stage, not per hire. Aggregate numbers hide stage-specific problems, and stages are where the fixes live.
  4. Never train on your own past decisions alone. Accuracy against historical choices is a measure of conformity.
  5. Version everything. If you can't say which rubric a candidate was scored under, you can't explain the decision six months later — and you will be asked.

    None of this makes hiring neutral. Nothing makes hiring neutral. It does make the judgements explicit, which means they can be argued with, corrected and defended — and that is a meaningfully different thing from a system that is quietly confident.
KEEP READING

Related Articles

Dive deeper into the ideas, research and playbook shaping smarter, and fairer hiring, curated to keep you one step ahead.
GET IN TOUCH

Finally, A Hiring Platform That Adapts To Your Process Not The Other Way Around.

GET IN TOUCH

Finally, A Hiring Platform That Adapts To Your Process Not The Other Way Around.

GET IN TOUCH

Finally, A Hiring Platform That Adapts To Your Process Not The Other Way Around.

GET IN TOUCH

Finally, A Hiring Platform That Adapts To Your Process Not The Other Way Around.

Create a free website with Framer, the website builder loved by startups, designers and agencies.