Skip to content

How to choose a hiring assessment tool

By Sergio Gualda · 11 min read · Updated

The choice is decided by two numbers before any feature list matters: how many different roles you hire for, and what a wrong hire costs you. Wide and cheap points to a test library; narrow and expensive points to a purpose-built instrument. Everything vendors compete on — question formats, integrations, dashboards — only starts to matter once those two numbers have already eliminated most of the market.

The two numbers that decide it

Every buying guide in this category opens with a feature comparison, which is why they are all useless. Every tool has every feature. What separates them is which problem they were built to solve, and that maps onto two things about you.

First: how many different kinds of role do you hire for? A company hiring accountants, warehouse staff, support agents and two engineers needs coverage. A company hiring product people needs depth. No tool is good at both, and the ones claiming to be are good at coverage and calling it depth.

Second: what does a wrong hire cost? At high volume the answer is low — you replace the person and the process absorbs it. On a five-person team hiring its second designer, the answer is a year, and possibly the shape of the whole product.

Wide and cheap: buy coverage, optimise for speed and price per candidate. Narrow and expensive: buy depth, and stop caring about price per candidate entirely — at three hires a year the difference between $20 and $200 per assessment is not a real number next to the cost of getting it wrong.

The five categories that actually exist

Vendors do not describe themselves this way, because every one of them wants to be all five. They are not. Named examples are illustrative, not endorsements — and one of these categories is ours.

  • Test libraries. Hundreds or thousands of ready-made tests across every function — TestGorilla, Testlify, Adaface. Right answer for breadth. You assemble an assessment per role from stock parts, and the score is a property of the candidate rather than of your job.
  • Work-sample platforms. The candidate produces real work and it gets graded, increasingly by AI — Vervoe is the clearest example. Strongest where the job has an observable output in an hour, which is true of far fewer roles than people assume.
  • Psychometric instruments. Personality and cognitive ability, grounded in decades of published research — Alva Labs, and the enterprise incumbents. Genuinely the most validated things on this list, and they measure the person rather than the fit with one specific job.
  • Game-based assessments. Short games that read cognitive and behavioural signal from how someone plays — Equalture. Excellent completion rates, deliberately role-agnostic, and they need enough hiring history to build a model from.
  • Enterprise hiring platforms. Video interviewing, AI screening, ATS integration, everything — HireVue and its peers. Solving throughput across an organisation, which is a different problem from choosing between four people.

The sixth category, and why it barely exists

Purpose-built instruments for one narrow band of roles. There are few of them, and the reason is economic rather than technical: a tool covering six roles can only be sold to companies hiring those six roles, which is a fraction of the market a library addresses.

What that constraint buys is the thing a library structurally cannot do — score the same candidate differently for two different jobs, because the instrument knows what those jobs are. A stock product designer test cannot know that your role needs someone who will fight about scope while the last one needed someone who ships what is agreed.

This is the category uxerhub is in, so read the previous paragraph accordingly. The honest version: if you hire across more than one function, you will end up buying a library too, and the narrow instrument is a supplement rather than a replacement.

What a demo will not tell you

Four questions that change the answer, and that no sales call volunteers.

  • What has this been validated against, and on how many people? The only question that matters. A straight answer names a criterion and a sample size. A logo wall is not an answer, and neither is «our customers report». Ask us this too — ours is that scoring is deterministic and the instrument is not yet calibrated against hiring outcomes.
  • What happens when a candidate contests a result? If a language model produced the score, explaining it is genuinely hard and the vendor knows it. Ask what the human review process is and who runs it.
  • How much unpaid candidate time does this cost, multiplied by everyone you assess? A ninety-minute assessment is not more rigorous than a fifteen-minute one. It is longer, and the people who decline are the ones with competing offers.
  • What does it cost to run one real candidate this week, without a call? A free trial that requires a demo is a lead form. This question also tells you what the company thinks of your time.

The costs that never appear on the pricing page

Three, and together they usually exceed the licence.

Candidate time, which you pay in declines rather than in money. It does not show up in your funnel — you see a shortlist and assume it represents the market, when it partly represents who had a free evening.

Review time. A tool producing rich output that someone has to interpret costs you an hour per candidate in a hiring manager's week, and hiring managers are the most expensive reviewers in the company. A tool producing a recommendation costs five minutes. This difference is larger than any price difference in the category.

Setup, which is where most implementations die. If configuring a role takes an afternoon, it gets done once, badly, and then reused for every subsequent opening regardless of what those openings need.

Where each category breaks

Nothing here is a criticism of a vendor. Every one of these is the direct consequence of what that category optimised for, which means it is structural and cannot be fixed with a feature.

Libraries break when a title covers jobs that need different people, because the test cannot know which job it is scoring for. Work-sample platforms break when the job's real output takes weeks rather than an hour, and when the artefact can be generated. Psychometrics break when a stable trait gets used to answer a question about one specific job. Game-based assessments break below the hiring volume needed to model your own performers. Enterprise platforms break on small teams, where the procurement cycle costs more than the hire.

And narrow instruments break the moment your next opening is not one of the roles they cover — which, for most companies, is soon.

A shortcut, if you want one

Answer these three and the field collapses to one or two options.

  • Do you hire for more than three distinct functions? If yes, you need a library, and the rest of this page is a supplement to that decision rather than an alternative to it.
  • Does the role have an output someone could produce and you could evaluate in under an hour? If yes, a work sample beats everything else. If no — and for most product roles the answer is no — you are choosing between measuring traits and measuring decisions.
  • Do two jobs with the same title in your company need different people? If yes, you need something that knows which job it is scoring for. Almost nothing does.

Frequently asked

What is the best hiring assessment tool?
There is no answer to that question as asked, and any page giving you one is ranking vendors it has a relationship with. The useful version is narrower: best for how many role types, at what cost of a wrong hire, and for roles with verifiable answers or roles without them. Those three make the choice nearly automatic.
Are free hiring assessment tools worth using?
Free tiers are generally enough to find out whether a category fits your problem, which is the expensive thing to get wrong. What they will not tell you is whether the instrument predicts anything — that requires asking about validation, and the answer costs nothing either.
How many assessments should a candidate do?
One, in almost every case. Stacking a personality test, a cognitive test and a work sample produces three hours of unpaid work before anyone has spoken to the person, and the marginal information from the third instrument is close to zero. Pick the one that answers the question you actually have.
Do assessment tools reduce bias in hiring?
They move it somewhere visible. Deciding in advance what the role needs and scoring everyone identically makes the judgment explicit and arguable rather than arriving as a feeling after a conversation. That is a genuine improvement and it is not the same as removing bias — someone still chose the criteria, and an instrument can carry bias in its own construction.