A Practical Checklist for Vetting Any AI Feature Your HR Vendor Just Shipped
Your ATS pushed an update with a new 'AI assistant' toggle — here's how to evaluate it before you turn it on for real candidates.
The pattern is familiar by now. You log into Greenhouse, iCIMS, Workday, SmartRecruiters, or Paradox on a Tuesday and there's a new panel: AI-generated interview questions, an "assistant" that drafts candidate rejections, a summarizer that condenses a resume into three bullets, a score you didn't ask for. It shipped in a release note. Nobody asked whether you wanted it, and the toggle may already be on.
These features are often genuinely useful. But "it came from a vendor we already trust" is not the same as "it's safe to point at live candidates." Here's the checklist I walk HR teams through before they flip one of these on for real.
1. What does it actually do — and does it decide anything?
Separate features into two buckets. Assistive features help a human do their own work faster: draft this outreach email, summarize this call, suggest interview questions. Decisional features influence who advances: score this candidate, rank this pool, recommend a shortlist.
The bar is completely different. An assistive drafting tool that writes a mediocre email costs you a rewrite. A decisional tool that quietly filters people carries legal and ethical weight and may fall under laws like NYC's Local Law 144 or the EU AI Act's high-risk category. Get precise about which bucket you're in, because vendors love to describe decisional features in assistive language. "Helps you prioritize" often means "ranks candidates."
2. What's the model, and where does the data go?
Ask three concrete questions and get answers in writing:
- What model powers this, and who hosts it? Is it running in the vendor's environment, or being passed to a third-party model provider? Both can be fine; you just need to know.
- Is our candidate data used to train anyone's model? The answer you generally want is no — that your data is processed to give you output and not retained for training. Get it in the contract or DPA, not a sales call.
- Does this change our data-processing footprint? A new sub-processor or a new region where data is handled can matter for GDPR and your existing candidate privacy notices.
3. Ask specifically about bias and testing
For anything decisional, ask what the vendor has done to test for disparate impact, and ask for documentation. Watch how they respond. A serious vendor has an answer — a bias audit, fairness testing, a methodology. A vendor that gets defensive, waves at "our AI is unbiased," or points only to the fact that the model "can't see race" is telling you something. Every model can learn proxies. "It doesn't see the protected attribute" is not a bias defense; it's a misunderstanding.
4. Can you see the reasoning, or just the output?
A score with no explanation is hard to defend and hard to trust. Prefer features that show their work: which parts of a resume drove a summary, why a candidate was flagged, what the recommendation is based on. If a candidate or a regulator asks "why was I screened out," "the software said so" is not an answer you want to give. Explainability isn't a nice-to-have here; it's what makes the feature auditable.
5. Who is accountable when it's wrong — and can you turn it off?
Two practical questions:
- Can you disable it? A per-recruiter and org-wide off switch should exist. Features that can only be "configured" but not fully turned off are a governance headache.
- Is the human in the loop real or theatrical? If the workflow technically requires a human click but the design nudges everyone to accept the AI's suggestion, that's a rubber stamp. Real oversight means the human can and sometimes does override — and you can show they did.
6. Test it on data where you already know the answer
This is the step teams skip, and it's the most valuable one. Before you trust a new feature, run it against a set of past cases where you know the outcome.
- For a summarizer: feed it 20 resumes you know well and read what it produced. Does it drop things that mattered? Does it emphasize the wrong things? Does it hallucinate credentials that aren't there? (This happens. I've seen a summarizer confidently invent a degree.)
- For a screener or score: run it against a batch of past applicants, including people you hired and people you passed on who turned out to be strong. Does it rank the people you know were good near the top? Where does it disagree with your judgment, and is it right or wrong when it does?
- For a drafting tool: generate a dozen outputs and check tone, accuracy, and whether it ever says something you'd never put in writing.
An afternoon of this tells you more than any vendor deck. You're not looking for perfection — you're calibrating how much to trust it and where it fails.
7. Document the decision
Once you've vetted it, write down what you found, who signed off, and on what date. When you turned it on, what testing you did, what the vendor told you about bias and data. This record costs ten minutes and is exactly what you'll want if a candidate complaint or a regulator's question lands eighteen months later. It also forces the decision to be a decision, rather than a toggle someone left on by default.
The one-line version
Before you point a new AI feature at real candidates, know what it decides, where the data goes, how it was tested for bias, whether you can see its reasoning, whether you can turn it off, and how it performs on cases you already understand — then write down that you checked. New features are guilty until tested, even from vendors you like.
This is general guidance for HR practitioners, not legal advice. Whether a specific feature triggers legal obligations depends on your jurisdiction and facts — involve qualified counsel and your privacy team before deploying decisional AI in hiring.
Put this into practice
Paste any text to estimate how many tokens it uses, and see what that text would cost to send to each major model.
Open the Token Estimator →A note on shelf life. AI products change fast. This guide deliberately focuses on the parts that stay true — how to judge a tool, what the trade-offs are — rather than ranking products that will have changed by the time you read it. Prices and feature claims should always be checked against the provider before you rely on them.