Case study · Recruitment platform (NDA)
Rebuilding a candidate search engine, settled by a benchmark
A recruiter describes in a plain sentence who they are looking for, and the system searches a base of a few thousand CVs and builds a ranking. This case study is about how I rebuilt that engine — and how I know the new version is better, not merely different.
The short version: the logs the search had been collecting for weeks became a benchmark that mirrors real traffic; the AI judge was calibrated like a measuring instrument; the new engine was built in isolation from production. Across the 50 benchmark postings it puts its list closer to the independent evaluation in 22 cases — the baseline in 2.
Production wrote the sequel. The new version replaced the old one, and I computed the post-deployment balance on 1,214 rows of real search logs. The biggest gain is not in the candidates' scores — it is in four classes of defect the client used to see on the list, and no longer does. Since then, quality telemetry computes the ranking's correlation with the judge for every real search: in September 2026 the mean Spearman for the reranked version is 0.72 across 108 searches — the level the benchmark predicted.
01Before the benchmark: telemetry and the first measurements
The search got telemetry first: every run records the full scoring components, the retrieval conditions and the rejection funnel — who dropped out, at which gate, and why. Without it, every conversation about quality ended in impressions; with it, those same logs later became the raw material the benchmark was built from.
The first round of measurement showed the problem was the data, not the model: the same skill existed in the database in over fifteen hundred spellings, and my own cross-encoder reranking R&D ended in a documented decision not to ship — it read the same text the embedding model did and hit the same ceiling, and the written-down conclusion was that the next leap would only come from LLM reranking. After the data was fixed, ranking agreement with the independent evaluation rose from 2% to 52%. This page is about the next step — that conclusion made real included.
02A benchmark from real traffic, annotated by hand
The collected logs yielded a profile of real queries: which roles, cities, tenure thresholds and language requirements recruiters actually ask for. The benchmark is 50 model postings mirroring that distribution — not invented cases, but a mirror of the traffic against a base of a few thousand CVs.
Each posting's expectations are written down by hand: what it requires, what it doesn't, which conditions are hard. By hand, because a benchmark whose expectations come from the code under test only checks the code's agreement with itself. A separate group of measurement postings differs from the baseline by exactly one condition — a tenure threshold, a salary range, an industry — so a difference in the result is the price of that one condition, not a blend of five causes at once.
03An AI judge, calibrated like an instrument
The measure of quality is the full AI report — an independent second opinion that never sees the system's own scoring, because a judge who knows the score stops being an independent measure. Before trusting its verdicts I measured the instrument itself: repeatability under deterministic shuffling of the order candidates are presented in (models favour what they read first), and a practical agreement ceiling of about 80% — that is what the same judge produces when run twice on the same list. One hundred percent agreement does not exist for any approach.
Calibration produced two lessons more general than this project. First: a general rule in the prompt — “don't score what the posting didn't ask for” — was broken in every third case, while the same rule given as an explicit computed list was respected one hundred percent of the time. Appealing doesn't work; enumerating does. Second: prompt variants, including enriching the judge with the full employment history, were measured too — and rejected, with the numbers and the reason written down so nobody rediscovers the same thing twice.
04A second engine — in isolation, in three phases
The new pipeline was built next to the production one, separated at the file level rather than behind configuration flags — users worked on the old version the whole time, and the production path stayed untouched to the byte.
Phase one narrows the pool with the posting's hard conditions already at retrieval from the vector database. Phase two is a set of reduction gates, each answering a single question; tenure is counted calendar-wise and within the query's scope — twenty years in retail plus two in accounting is two years when the search is for an accountant — and the funnel records the complete set of rejection reasons, with a counter for “how many candidates come back if this condition is relaxed”. Phase three scores only the dimensions the posting actually sets and the filters haven't already settled: the share of 97+ scores that differentiated nothing fell from half of all recorded results to zero.
At the end of the pipeline stands a light LLM reranker — the conclusion from the rejected cross-encoder, made real. It orders the finished shortlist but has no right to reject or add anyone, and any deviation of the response from a full permutation means falling back to the scoring order, with a trace in the log. This step, too, went in after a measurement, not on faith: in the pilot its ordering agreed with the verdict of a stronger model 66% of the time, the scoring alone 39% — and it is chiefly this step that accounts for the ordering agreement with the AI report in the results below.
05Weights from measurement, not intuition
Which components actually predict the judge's verdict? I computed the correlations within a single posting — where the ranking is made; computed across searches they gave the opposite conclusions, and that was the earlier methodological mistake. The result: skill matching turned out to be a random signal — the sign of the correlation flipped from posting to posting — while carrying about two thirds of the score. The only consistent predictor was vector similarity.
The new weights sit in the middle of a wide plateau rather than on a spike fitted to noise, and validation on a held-out posting confirmed the direction: agreement with the judge rose from 15% to 41%, top-five coverage from 58% to 78%.
06Tests without calling the model
In this pipeline the model's answer turns into hard filters, so edge cases can't be studied against the live API: it costs money, it isn't deterministic, and it doesn't stay in the repository. The fake model returns raw text, not a ready-made object — because half of the real defects live in the text: markdown fences before the JSON, an introductory sentence, a response cut off at the token limit.
Tests written from the business requirements before the code found real bugs — a declared B2 level read as fluency because the more general pattern was checked first, and a city vanishing from the conditions through a type mistake in the model's answer, after which the pool covered all of Poland and the result looked normal while answering a different question. The suite grew past 440 offline tests; the prompt contract stands separately, because it is the only one that costs money.
07The result: 50 postings, both approaches, the same judge
Each of the 50 benchmark postings went to both versions, and each list was scored by the full AI report — with each approach's funnel honestly described, so neither gets points for criteria its own filters had already settled. The whole test cost about 4 USD.
Ordering agreement with the AI report
First on the list = best per the AI report
practical ceiling ≈ 80%
Top-three overlap
Position of the best candidate per the AI report
lower = better
new version2.0
baseline3.1
22 : 2
postings where a given version puts its list closer to the AI report (new : baseline; plus 7 ties among 31 decidable comparisons)
How both versions arrive at a list of ten people
new version
baseline
The new version's hard filters did not shrink the lists — there are more full 10/10 lists (30 vs 27 out of 50). The difference lies earlier: a smaller, better-aimed pool already at retrieval, so the filters have less to cut and more candidates reach scoring. About 90% of the new version's filtering comes from two posting conditions: required languages and the tenure threshold.
The price of this quality is computed, not glossed over: mean search time 22.6 s vs 11.5 s — the difference is mostly the model's ordering step — and a per-query cost of about 0.008 USD vs 0.003 USD.
08Deployment: the new version takes over the traffic
The same result across three measurements
The same 50 postings: the baseline and two stages of the new engine.
| measurement | list composition | ordering |
|---|---|---|
| baseline | 59.0 | 0.35 |
| new — gates and quality threshold | 64.3 | 0.69 |
| new — after city geocoding | 65.4 | 0.69 |
Ordering settled at 0.69–0.73 across three consecutive measurements, and that is the important sentence here: the benchmark result was not the luck of a single run. It is also a ceiling — the filters will not move it any further, and the next leap would have to come from the ranking. List composition gains some six points, but the last step sits inside the judge's noise, so what counts is the direction, not the third digit.
How the switch itself was done
A benchmark settles the argument, but what ships is code, not a benchmark. Between the test and the switch came one more step: a manual review of six postings by a human tester. The sharpest case: on a financial-controller posting the old version ordered the list inversely to the human's ratings, and the new one in line with them.
The switch itself I did the harder way round: the new engine's code moved into the production module, rather than client traffic moving onto the research module. Rerouting the traffic would have produced a pipeline I had never measured end to end — a new list scored by an old prompt. The entire production casing stayed untouched: logging, personal-data anonymisation, the response contract. The engine changes, the functionality does not.
Before the deployment, not after: proof the copy was faithful (the differences against the research version are only the intended edits), an equivalence test of results on eight postings, an anonymisation check on the client path, and an exit route described up front — no database migration, so rolling back recomputes nothing. Every stored search now carries an engine-version marker, so a year from now the rows from before the switch can still be told apart from the ones after it.
Four classes of defect that left production
Ordering can be measured without noise; an average score cannot — so the second piece of evidence is not points but four classes of defect that vanished from the list the client sees. I counted the scale of each in the logs of real searches before the fix went in.
Industry guessed from the employer's description and treated as a hard requirement
- in the logs
- It cut more than half the pool in 9% of client queries and was the main cause of 8 of the 20 empty lists.
- after the change
- In the tester's review it never set an industry at all; both empty lists filled up with candidates.
No quality threshold — anyone past the filters made the list
- in the logs
- 6% of the positions shown to clients scored below 20 out of 100; the tester was handed people scoring 1, 4 and 9.
- after the change
- An entry threshold; in the first 33 searches after deployment it cut 35 candidates.
No tenure gate — missing experience cost points but never eliminated
- in the logs
- On a posting requiring five years the list showed candidates with six months of experience.
- after the change
- The gate removed 9 of 13 people — and it was the remaining four the judge scored highest, 70 against 49 for the full nine.
The remote-work filter cut both ways
- in the logs
- An on-site posting removed everyone open to remote work — that is 45% of the base.
- after the change
- A one-way filter; two people came back whom the judge scored above anyone on the pre-change list.
Reproducibility, which no benchmark had asked for
47%
of repeated searches returned an identical order — despite zero temperature and a fixed seed
This came out of a measurement on the logs after the deployment. A user clicking “search” a second time got a reshuffled list, and worse, my own comparisons between the two versions were partly measuring model noise — so the problem touched not just the product but the method. The ordering is now remembered under a key computed from the whole material the model saw: a repeated search comes back identical, and a dozen-odd seconds faster into the bargain.
The first hours in production
33
searches in the 47 minutes after the switch
0
failures of the ordering step
35
candidates cut by the quality threshold
1
empty list out of 33 searches
The panel, the site and the research environment returned an identical set and an identical order of fifteen candidates for the same posting. Honestly: 31 of those 33 searches were test traffic, so the first day's logs confirm that it runs, not that it is better — the quality claim rests on the pre-deployment measurements.
There is exactly one regression and it is computed too: median search time rises from 8.4 s to 27.9 s, of which 13.3 s is the model's ordering step alone, at a cost of about one US cent per search. The bill has a non-obvious shape — for a few dozen tokens of output the model burns some 2,800 “thinking” tokens billed as output, which is 85% of the call's price.
An honest caveat
The reference is a score handed down by a model, and that measure has two flaws, both written into the deployment report. It is not stable in absolute terms — it gave the same seven candidates 90 in one run and 66 in another — so differences of a few points in the averages are the judge's noise, not the pipeline's progress. And it is anchored: the prompt shows the judge our own score before it issues its own, so every agreement figure is an upper bound. That is why ordering and structural differences count here, not the average score. Two things stay open and both are counted: surplus experience still costs nothing, so a senior lands on a junior posting, and the distribution of list lengths is bimodal — either a full list or almost nothing.