How much do you actually trust an AI-written job-match score?
by•
Working on scoring candidates against job postings with AI, and I keep going back and forth on this: a bare percentage feels untrustworthy no matter how good the model is, but a full paragraph of reasoning is too slow to skim across 50 profiles.
Landed on a point-by-point checklist (what's met, what's missing) as a middle ground, but curious how others handle this. If a tool told you "73% match," would you trust it without seeing why? Where's the line between useful signal and noise for you?
66 views
Replies
A percentage without context feels like a credit sore for a human 😂. I’d much rather see the app the top3 things that drove the score.
I ran 500+ assessments and 200+ tech interviews, and honestly humans don't trust their own scores either. That's why we always made people write the reasoning first and the number second. Number first and you spend the rest
of the review defending it.
So for me the question isn't "do I trust 73%". It's what the number is for. If it sorts 50 profiles so I read the right 10 first, I trust it fine, being roughly right is enough for ranking. If it decides anything on its own, no amount of reasoning text would save it.
Your checklist is the right call, but I'd flip the emphasis: show the misses, not the matches. When I skim 50 candidates I'm not looking for reasons to say yes, I'm looking for the one gap that makes it a no. Three red items I can scan in a second beat a balanced summary I have to read.
The thing I'd actually want: let me disagree once and have it stick. Mark "missing Kubernetes" as irrelevant for this role and never see it counted again. Trust doesn't come from better explanations, it comes from the tool losing an argument.
The score should help me decide where to look, not decide for me. I would find a short evidence based breakdown much more useful than a percentage, especially when one missing requirement can completely change a candidates suitability.
A single 73% hides too much. I’d separate hard requirements from soft-fit signals, then use the percentage only for ranking. Someone missing one critical requirement shouldn’t look equivalent to someone with a few minor gaps. That would make the score much easier to trust at a glance.
ClawTeams
Siarhei's point about writing the reasoning before the number is the one I'd build around, and I'd push it one step further.
The deeper problem with 73% is that you can never check it. For a number to earn trust it has to be calibrated: out of everyone scored 73, roughly 73% should work out. But you only ever learn the outcome for the people you hired, and those are the high scores. The whole bottom of the distribution is unobservable forever, so the number can't be validated even in principle. It's not that the model is wrong, it's that "correct" isn't defined for it.
A rank makes a much weaker claim: this one before that one. You can actually be right about that, and it's all you need if the job is deciding who to read first.
The other cost is that showing a number moves the reader. Once someone sees 73 they read the profile looking for 73. If you show the checklist first and put the score behind a click, you get the sorting without the anchor.
And the overrides are your only real eval set. Every time someone reads a low score and hires anyway, that's ground truth you can't get any other way. Worth logging from day one.
Tony's calibration point sharpens Siarhei's argument: you only ever observe outcomes for the people you hired, so the number can't really be validated even in principle. A rank sidesteps that, it only claims this one before that one, which you can actually check. The checklist is still right for what a person reads, just keep the score behind a click so it doesn't anchor the reading before anyone's looked at specifics. Logging every override where a low score got hired anyway is probably the most useful data in this thread.