We scored 40,000 applications twice. Here's the drift.

Nadia Okonkwo
HR Specialist
Latest update
14 min read

IN THIS ARTICLE
The Four Fields Problem
Proxies that survive annoymization
When the rubric itself is the bias
Feedback loops, and why they hide
What monitoring actually catches
A short checklist
10 min remaining
Almost every recruiting team that tests AI scoring does the same first thing: they score the same applications twice and compare the results. It feels like the hard part is done. The model can now apply the same criteria consistently, so the scores should remain stable from one run to the next. But small changes in scoring can reveal drift — and drift is exactly what repeated evaluation is built to find.
Each of those carries signal about the candidate. Several of them also carry signal about the scoring process. The model cannot tell the difference, because from its position there isn't one — it sees application text that correlates with outcomes across repeated runs and weights it accordingly.
The four fields problem
Consider what remains after a second scoring pass: candidate scores, ranking changes, pass rates, flagged applications, score differences, reviewer notes, model outputs, and the patterns that shift between runs.
Each of those carries signal about the candidate. Several of them also carry signal about the scoring process. The model cannot tell the difference, because from its position there isn't one — it sees application text that correlates with outcomes across repeated runs and weights it accordingly.
The test we use internally
Take a repeated application score from your pipeline and hand it to two colleagues. Ask them to explain why the candidate received that score and whether they would rank the application differently. If your colleagues reach different conclusions, so can a model — and repeated runs can reveal the pattern through score drift.
Proxies that survive anonymisation
These are the patterns we see cause measurable score differences most often, roughly in order of how much impact they have relative to how easy they are to overlook:
Employment gaps. A fourteen-month gap can create a measurable score difference between repeated evaluations. Models trained on “continuous tenure predicts performance” may penalise it heavily, and the penalty can shift depending on how the application is scored.
Postcode and commute distance. Where someone lives can create measurable differences in repeated scoring. If your rubric rewards proximity to the office, you have built a geographic filter that can shift scores between otherwise similar applications.
Prestige signals in prose. Candidates coached on CV writing use a specific register — action verbs, quantified outcomes, tight parallel structure. That register can shift how applications are scored between runs. Reward the writing and you may reward the coaching rather than the candidate’s actual experience.
Non-linear careers. Someone who moved into tech at 34 has a shorter tech history than someone who started at 22, and repeated scoring can read that as less experience rather than a later start.
Volunteer and interest lines. Community groups, sports, cultural associations. Almost nobody scores these deliberately, but repeated scoring can still pick up signals from them and shift how an application is evaluated.
When the rubric itself is the bias
The uncomfortable finding from our own scoring runs is that most measurable drift does not originate in the model. It originates in the criteria the model was told to apply.
A rubric that requires “8+ years in a similar role” can produce different pass rates across repeated scoring runs. A rubric that heavily weights “startup experience” can favour applications with that specific background. A rubric that treats a named certification as mandatory can filter out candidates trained elsewhere. None of this requires a biased model. A perfectly calibrated model applying these criteria produces the same disparity the criteria encode, faster and more consistently than a human would.
This is why we spend more engineering effort on how criteria are written than on the scoring model itself. It’s also why every rubric is versioned: when a pass-through rate shifts, the first question is always which criteria changed, not whether the model drifted.
Feedback loops, and why they hide
The slowest failure is the one that compounds. If a scoring system is tuned on your past decisions, it learns what your organisation has historically rewarded — including the patterns you're trying to change. Its scores then shape the next round of evaluations, which becomes the next set of data.
Two years in, the system looks highly accurate. It predicts your hiring decisions with impressive consistency. That accuracy is the problem: it is measuring agreement across scoring runs, not whether the process produces better outcomes.

Stage-level pass-through monitoring. A score difference that appears at one stage and disappears at the next usually points to a criterion, not the model.
What monitoring actually catches
The two criteria they changed: continuous employment history became “relevant experience, in any arrangement”, and a specific cloud certification became “that certification or demonstrable equivalent”. The screening rate moved from 0.71 to 0.94, and seventeen more candidates reached interview — from the same applicant pool.
The point isn’t that rubric edits fix everything. It’s that the drift was visible at a specific stage, in week two, while the roles were still open. Annual reporting would have surfaced it months later, attached to hiring decisions that had already been made.
A short checklist
Audit the rubric before the model. Read every criterion and ask which candidates it excludes that the job doesn’t actually require.
Test for repeatability. If two scoring runs produce different results from the same application, treat the scoring process as unstable.
Monitor per stage, not per application. Aggregate numbers can hide stage-specific drift, and stages are where the fixes are found.
Never train on your own past decisions alone. Accuracy against historical choices is a measure of conformity, not whether those choices produced better outcomes.
Version everything. If you can’t say which rubric a candidate was scored under, you can’t explain the decision six months later — and you will be asked.
None of this makes scoring stable. Nothing makes scoring stable. It does make the differences visible, which means they can be reviewed, corrected and explained — and that is a meaningfully different thing from a system that is quietly confident.







