Skip to content

Personality tests for hiring designers

By Sergio Gualda · 10 min read · Updated

Self-report personality instruments measure what a candidate believes about themselves, or what they think you want to hear — which in a hiring context is the same thing. They predict job performance far more weakly than work samples, create legal exposure when used as a selection gate, and are visibly resented by experienced candidates.

The desirable answer is always obvious

Nobody applying for a collaborative role selects “I prefer to work alone and rarely seek feedback”. Self-report instruments assume good-faith introspection, and a hiring process is the single situation where that assumption is least safe.

You are not measuring the trait. You are measuring the candidate's model of what you want, which is a real skill but not the one on the job description.

Selection and development are different uses

Used after hiring, to help a team understand how its members prefer to work, the better instruments are genuinely useful and largely harmless. The failure is using them as a gate.

As a gate they carry legal exposure that varies by jurisdiction but is never zero — particularly where a trait correlates with a protected characteristic, or where the instrument has never been validated for the specific role.

What senior candidates read into it

Experienced designers recognise a personality questionnaire in a hiring flow immediately, and a meaningful proportion of them withdraw. Not on principle in the abstract, but because it signals that the company cannot assess the work and is looking for a proxy.

That signal is often correct, which is why it is worth taking seriously.

Measure the behaviour instead

Everything a personality test is trying to get at can be observed directly, if you put the person in the situation and ask for the artefact.

Instead of asking how someone handles disagreement, put a stakeholder in front of them who has decided the answer and will not move, and ask them to write the reply. Instead of asking whether they take feedback well, criticise their proposal halfway through the exercise and see what they do with it.

You end up with words you can read and judge, rather than a trait score you have to take on faith.

The five interpersonal signals worth capturing

These are observable inside a work task, and each produces a real artefact.

  • Explaining a decision to someone outside the discipline, without jargon.
  • Handling disagreement with a senior stakeholder who will not move.
  • Taking hard criticism mid-task — revise, defend, or collapse.
  • Working with an engineer who says the approach is not possible.
  • Self-awareness: naming the weakest part of their own proposal.

What the research actually says, in both directions

The case against personality testing in hiring is usually overstated by people selling something else, so here is the fair version.

Personality measures do predict job performance — modestly, consistently, and better than an unstructured interview. Conscientiousness in particular shows a small positive relationship with performance across most job families, and it has held up across decades of study. Anyone telling you these instruments measure nothing is wrong.

The problem is not validity, it is application. A small average effect across thousands of people is a real finding and an almost useless basis for deciding between two specific candidates, because the variation between individuals swamps it. It is the difference between a signal that shows up in a population and a signal you can act on in a room.

And the instruments used in most hiring processes are not the ones the research validated. A four-letter type indicator is not a Big Five inventory, does not have its psychometric properties, and was never designed for selection — its own publisher says so.

The specific failure mode in design hiring

Design teams reach for personality testing at a predictable moment: when the shortlist is technically strong and someone says «but will they fit here». That is a real question and this is the wrong instrument for it.

What the team usually means by fit is one of three things — will they cope with our level of ambiguity, will they push back or go quiet, will they survive the review culture. All three are situational. They depend on the environment as much as the person, and the same designer can be excellent in one and miserable in the next.

A trait score cannot answer a situational question. What it does instead is give the discomfort a number, and a number is very hard to argue with — which is how a team ends up rejecting a strong candidate for being «too low on agreeableness» when what actually happened was that one interviewer found them abrupt.

If you are going to use one anyway, five rules

Sometimes the decision is not yours, or the instrument is already in the process. These rules keep it from doing damage.

  • Never use it to screen out. Use it after the decision to hire, to inform how you manage and onboard the person. That is the use these instruments were built for and where they hold up.
  • Never let it break a tie. If two candidates are close on the things that matter, a trait difference is noise, and using it to decide feels rigorous while being arbitrary.
  • Decide what the profile means before you see anyone. Written down, in advance. Otherwise the profile gets interpreted to support the preference that already exists.
  • Use an instrument with published psychometric properties, and read them. If the vendor cannot tell you the reliability coefficients or what it was validated against, it is a personality quiz with a sales team.
  • Tell the candidate what it is for and what you will do with the result. If that sentence is uncomfortable to write, that is information about the use, not about the writing.

What to measure instead, and how

Everything a team actually wants from a personality test is available from behaviour, and behaviour is both more predictive and easier to defend.

Ambiguity tolerance: give them a brief with a hole in it. Do they ask, assume and flag the assumption, or build something and hope? All three happen, they are visibly different, and no self-report question distinguishes them.

Conflict behaviour: ask about a decision they lost. Not how they handle conflict in general — everyone answers that identically and well — but a specific occasion, what they did afterwards, and whether the work continued. The general question measures composure under interviewing; the specific one measures what happened.

Feedback response: give real critique on something they presented, in the interview, and watch. This is uncomfortable and it is the single most informative thirty seconds in most design hiring processes, because it is the one moment the candidate cannot have rehearsed.

Frequently asked

Is MBTI acceptable for hiring?
No serious practitioner recommends MBTI for selection. Its test-retest reliability is poor enough that a meaningful share of people get a different type weeks apart, and it was never designed as a selection instrument.
What about cognitive ability tests?
Different question, and a genuinely different evidence base — general cognitive ability is among the better-supported predictors in the research. They come with their own fairness considerations, but they are not in the same category as self-report personality inventories.
Are personality tests legal to use in hiring?
In most jurisdictions yes, with conditions that vary and that tighten considerably where a result could be read as health-related information. In the EU, automated decisions that significantly affect someone carry specific obligations, including the right to human review. This is not legal advice — the point is that if a personality result contributes to rejections, that is a decision you need to be able to explain and defend, and many teams using these tools have never checked whether they can.
What about culture-fit assessments? Are those different?
Usually the same instrument with a friendlier label. The test is what happens to the result: if a low score removes someone from the process, it is a selection instrument regardless of what it is called, and it carries every problem described above.
Do situational judgment tests have the same problem?
They have a different one. An SJT measures a decision in a described situation rather than a trait, so it is tied to the job rather than to the person — but the desirable answer can still be guessable if the options are not built carefully. That is why a well-built one offers no obviously correct option: every choice has to be something a reasonable professional might do, or you are measuring who guessed what you wanted.