Skip to content
tal.co
· 8 min read

How to QA an AI Sourcing Agent Before It Reaches Hiring Managers

A step-by-step audit for checking accuracy, bias, and fit in AI-sourced candidate lists before hiring managers see them, with the exact metrics to track.

R
Ryan Nead
September 22, 2026

An AI sourcing agent will happily deliver a hundred candidates before your coffee is cold. That is the pitch, and on a good day it is true. The problem is the day it is not, when the list looks fine on the surface and quietly contains the wrong seniority, invented job titles, or a slate that would fail a basic adverse impact check. By the time a hiring manager notices, three interviews are on the calendar and trust in the tool is halfway out the door.

The fix is not to slow the agent down. It is to put a lightweight QA step between the agent's output and the hiring manager's inbox, one a recruiter can run in under an hour per requisition. This piece walks through how to score, sample, and audit an AI sourcing agent's list so a human is still the last set of eyes before anyone gets contacted.

So what does that QA loop actually look like in practice?

Define What "Good" Means Before the Agent Runs

Most AI sourcing failures start upstream. The agent is graded against a brief that was never precise enough to grade against. Before you evaluate a single profile, write down the exact criteria the list will be scored on, and make them binary or numeric wherever possible.

A workable rubric for a single requisition usually contains four buckets:

  • Must-haves the profile has to satisfy (specific credential, location radius, work authorization).
  • Strong signals that raise a candidate's score (relevant company size, domain, tenure pattern).
  • Disqualifiers that push a candidate off the list regardless of other strengths.
  • Ambiguity flags that require a human check rather than an auto-reject, such as gaps or non-linear paths.

This rubric is what the QA pass actually measures against. Without it, "the list looks bad" is a feeling, not a finding, and the agent's owner has nothing to tune. Teams that have written down their intake in this shape also tend to have fewer arguments with hiring managers about which resume red flags matter for this specific role, because the answer is on paper.

Sample the List, Do Not Read the Whole Thing

You do not have time to read 200 profiles per role, and you do not need to. Statistical sampling is how every other quality function in a mature org works, and sourcing QA is no different. A defensible sample for a single-role batch is 20 to 30 profiles, drawn in three slices:

  • Top of the ranked list (roughly 10 profiles). These are what the hiring manager will actually see first, so precision at the top matters most.
  • Middle of the distribution (roughly 10 profiles). This is where drift usually hides. If the top ten look great and the middle looks like a different job entirely, the model is over-fitting a narrow pattern.
  • A random tail draw (roughly 5 to 10 profiles). This catches the "why is this person even here" cases that reveal how loose the retrieval got.

Score each profile against the rubric on a simple pass/fail plus a one-line note. The output is a precision number you can track over time: of the sampled profiles, what percent cleared every must-have. Anything under about 70% on the top slice is a signal to pause the batch and re-brief the agent rather than push it to the hiring manager. This turns data into hiring decisions instead of vibes.

How to Sample a 100-Profile Batch
How to Sample a 100-Profile BatchTop of ranked list: 10 profiles; Middle of distribution: 10 profiles; Random tail draw: 8 profilesTop of ranked list36%Middle of distribution36%Random tail draw29%
A defensible QA sample pulls from three slices of the ranked list, not just the top. Illustrative: a visual comparison, not measured data.

Check for Hallucinated Fields, Not Just Bad Fits

Fit is the obvious failure mode. The subtler one is fabrication. Generic LLMs used in resume parsing frequently produce hallucinated outputs, including invented job titles or skills that were never in the source document. That means a sourcing agent's enriched profile can look richer than the underlying evidence supports, and a recruiter who trusts the enrichment ends up pitching a hiring manager on qualifications the candidate never claimed.

Two checks catch most of it. First, spot-verify five profiles per batch against the primary source (the candidate's public profile, portfolio, or resume) and confirm every claimed title, employer, and skill actually appears there. Second, watch for suspiciously clean matches. If every profile in a technical batch lists the exact same five frameworks in the exact order your JD lists them, the agent is likely echoing your brief back at you rather than reading the candidate.

Related failure: unjustified confidence on close calls. In one study of LLM-based resume screening, most models scored below 0.90 on discriminant validity, arbitrarily picking a winner in over 10% of equal-pair cases instead of abstaining. Your QA should reward the agent that says "I don't know" on ambiguous profiles more than the one that always ranks confidently.

A recruiter's desk with a checklist clipboard next to a laptop showing candidate tiles

Run a Simple Adverse Impact Screen on the Slate

Sourcing lists are a selection procedure in the eyes of federal regulators. EEOC guidance treats algorithmic decision-making tools as a "selection procedure" under Title VII, and the employer remains liable for disparate impact even when the tool is built by a third-party vendor. Vendor assurances are not a shield.

The practical standard to hold each batch against is the four-fifths rule. If the selection rate for a protected group is less than 80 percent of the rate for the highest-selected group, that is evidence of adverse impact. On a single requisition your numbers will usually be too small for that math to mean much statistically, so run it two ways:

  • Per-batch sanity check. Compare the demographic mix of the agent's output to the demographic mix of the qualified labor market you drew from. A wildly narrower slate is a flag to investigate before you send it on.
  • Rolling audit across batches. Aggregate 90 days of sourced-and-advanced candidates and compute impact ratios there. This is where recruitment algorithms drift as the candidate pool changes, so a quarterly cadence catches the model that looked balanced in Q1 and quietly narrowed in Q3.

Teams operating in New York City have an extra reason to keep this tidy: non-compliance with NYC Local Law 144 can result in civil penalties of $500 to $1,500 per violation, with each day treated as a separate violation.

Four-Fifths Rule, Applied to a Sourced Slate
Four-Fifths Rule, Applied to a Sourced SlateReference group (highest rate): 1 ratio; Group B: 0.9 ratio; Group C: 0.8 ratio; Group D: 0.6 ratioReference group(highest rate)1 ratioGroup B0.9 ratioGroup C0.8 ratioGroup D0.6 ratio
Selection rates for each group are compared to the highest group. Ratios under 0.80 flag investigation. Illustrative: a visual comparison, not measured data.

Grade the Agent, Then Feed the Grade Back

A QA pass that stops at "approved" or "rejected" wastes half its value. The other half is the feedback loop that makes next week's batch better. Keep a running scorecard per agent, per role family, with four numbers:

  • Top-slice precision. Percent of the top ten profiles that cleared every must-have.
  • Hallucination rate. Percent of spot-verified profiles that contained at least one fabricated field.
  • Recruiter override rate. Percent of the agent's ranked order that a human reshuffled before sending.
  • Downstream hit rate. Percent of sent profiles that advanced past the first hiring-manager screen.

These four numbers, tracked for even a month, will tell you more about an AI sourcing agent than any vendor deck. The downstream hit rate in particular is the one that matters to the business, because it is the one that maps to speed and quality in hiring at the same time.

Sourcing Agent Scorecard, Per Role Family
MetricWhat it measuresHealthy rangeAction if breached
Top-slice precisionMust-haves cleared in top 10Above 70%Re-brief the agent
Hallucination rateFabricated fields in spot checkUnder 5%Turn off enrichment
Recruiter override rateRank reshuffle before sendUnder 30%Retune ranking weights
Downstream hit rateAdvanced past HM screenAbove 40%Revisit intake with HM
Four numbers per agent per role family, refreshed weekly. Illustrative: a visual comparison, not measured data.

Keep a Human in the Loop Where It Actually Matters

Human-in-the-loop is a phrase that gets used loosely enough to mean nothing. In sourcing QA it means something specific: a named recruiter signs off on the batch before contact goes out, and that sign-off is recorded. Not every profile needs a full read. The signature is on the sample, the scorecard, and the adverse-impact screen.

The recruiter's job at that gate is judgment work the agent cannot do. Reading between the lines of a non-linear career. Deciding a two-year gap is a caregiving story, not a red flag. Noticing that the top five profiles all come from one competitor and the hiring manager will find that awkward. These are the calls that separate a sourcing function from a scraping function, and they are the reason the recruiter's role in AI hiring is expanding rather than shrinking.

None of this needs to be heavy. A good QA pass on a 100-profile batch takes 30 to 45 minutes once the rubric exists. Given that recruiters already spend roughly 13 hours per week sourcing candidates for a single role, spending under an hour to protect the quality of what leaves the building is a reasonable trade.

What Changes When You Actually Do This

Two things change quickly. Hiring managers stop pushing back on AI-sourced lists, because the lists they see have already been filtered by a person they trust. And the agent itself gets sharper, because a written scorecard is the input the vendor (or your own tuning process) needs to fix what is actually broken instead of guessing.

AI sourcing is a tool, not a verdict. The recruiters getting the most out of it are the ones who treat every batch as a draft that needs a QA pass before it becomes a decision. The rubric, the sample, the fabrication check, the four-fifths screen, the scorecard: five steps, an hour of work, and the difference between an agent that earns its keep and one that quietly burns down the hiring manager's confidence in the whole program.

Build a hiring system, not another inbox.