Skip to content
uxerhub

Situational judgement tests, and when they are the right instrument

By Sergio Gualda · 7 min read · Updated

A situational judgement test presents realistic work scenarios and asks what the candidate would do, scoring which action they reach for rather than what they know. It sits between a skills test, which needs a verifiable answer, and a personality questionnaire, which asks people to describe themselves — and it is used when the decisions that matter in a role have no correct answer. The format has a long history in selection research and it is also easy to build badly: the common failure is a scenario with one obviously right option, which measures nothing except whether the candidate can spot what the employer wants to hear.

What an SJT is, precisely

A scenario from the job, a set of possible responses, and a scoring key derived from what capable practitioners actually do. The candidate marks what they would do — sometimes what they would do first and what they would never do, which extracts more signal than a single choice.

The defining property is that the options are all plausible. A well-built SJT has no obviously correct answer, which is what stops it from measuring social desirability instead of judgment.

That property is also the whole difficulty of building one. Writing four defensible options for the same situation is much harder than writing one right answer and three wrong ones, and the shortcut is visible immediately to a candidate who knows the job.

How it differs from the neighbouring formats

Three comparisons that place it, and each explains when it is the wrong choice.

  • Versus a skills test — a skills test needs a checkable answer and is superior wherever one exists. If you want to know whether someone can write working SQL, use a SQL test. An SJT is for the decisions that have no key.
  • Versus a personality questionnaire — a questionnaire asks people to describe themselves, which invites them to describe the person they think the job wants. An SJT asks what they would do in a specific situation, which is harder to game because the desirable answer is not obvious. Personality instruments have the far stronger validation record.
  • Versus a work sample — a work sample is more direct evidence and generally the better predictor. It costs hours of unpaid candidate time, so it belongs on a shortlist, while an SJT is cheap enough to send to a whole pool.

Where situational judgement tests fail

Four failure modes, three of which are the test's fault.

  • Transparent options. If one response is obviously what the employer wants, you are measuring social awareness and nothing else. This is the most common failure and the easiest to spot: read the scenarios yourself and see whether you can pick the intended answer without knowing the job.
  • Scenarios that do not resemble the work. A generic «a colleague misses a deadline» item measures a generic disposition. The value comes from situations specific enough that a practitioner recognises them and an outsider cannot bluff them.
  • A key derived from opinion rather than practice. The scoring has to come from what capable people in that role actually do, and a vendor should be able to say whose judgement built the key.
  • Uncalibrated deployment. An SJT that has not been checked against hiring outcomes is a structured opinion. That is still better than an unstructured one, and it is not the same as a validated instrument, and anyone selling you one should say which they have.

How to tell a good one from a questionnaire in disguise

Ask for the actual items and read them. Vendors who will not show you scenarios before a purchase are telling you something.

  • Can you pick the intended answer without knowing the role? If yes, it is measuring social desirability.
  • Is every option something a competent person might genuinely do? If one is a straw man, the item is doing no work.
  • Do the situations name real constraints — a date, a stakeholder, a resource that just disappeared? Judgment only becomes visible when something is being traded away.
  • Does the score change depending on the role you are hiring for, or is it fixed? A stock key cannot know that your job needs someone who fights about scope while the last one needed someone who ships what is agreed.
  • What has it been validated against, and on how many people? Ask everyone, including us.

Where they fit in a hiring process

Near the top, after a basic screen and before anything expensive. The economics are the point: fifteen to thirty minutes is cheap enough to ask of everyone who clears your first filter, which means the comparison covers the whole pool rather than the three people who already looked promising on paper.

That is where the value concentrates, because the candidates who look unpromising on paper are precisely who a screening instrument exists to catch. An instrument only given to the shortlist cannot do that job by construction.

What it does not replace is the interview. A good SJT report should make the interview better by telling you what to ask, not shorter by telling you who to hire.

Where we fit, stated plainly

uxerhub is a situational judgement test, so this entire guide describes the category we sell in. Read it accordingly and apply the checklist above to us.

Our answers to it: twelve situations per role, written for six product roles specifically rather than generically; four plausible options per item with no intended answer; each situation names a constraint being traded away; and the score is weighted by up to three capabilities the hiring company names before sending the link, so it is not fixed across buyers.

And the one we fail: it is not calibrated. Scoring is deterministic arithmetic rather than a model, but we have not run enough hires through it to publish predictive validity, and by the standard set in the section above that makes it a structured opinion rather than a validated instrument. Personality and cognitive instruments from established vendors have evidence we do not.

Frequently asked

What does a situational judgement test measure?
Which action someone reaches for, and which they rule out, when faced with a realistic work scenario in which no option is clearly correct. It targets judgment rather than knowledge, which is why it is used for roles whose important decisions have no verifiable answer.
Are situational judgement tests accurate?
The format has a long standing in selection research, but accuracy is a property of the specific instrument rather than the category — a badly built SJT with transparent options measures almost nothing. Ask any vendor what their key was derived from and what it has been validated against, on how many people.
How is an SJT different from a personality test?
A personality questionnaire asks people to describe themselves, which invites them to describe who they think the employer wants. An SJT asks what they would do in a named situation, where the desirable answer is not obvious. Personality instruments have the stronger validation record; SJTs are harder to game item by item.
Can candidates cheat a situational judgement test with AI?
They can use one, and a well-built test contains it structurally rather than by detection: if a fluent, hedged, cover-everything answer scores in the middle by design, then the model's default register is not an advantage. The points have to sit in committing to one option under a constraint and naming what you give up.
How long should a situational judgement test be?
Short enough to send to everyone who clears your first screen — in practice under thirty minutes. Its economic advantage over a work sample is entirely in the cost to the candidate, and a ninety-minute SJT has given that advantage away while keeping the weaker evidence.