a situational judgment test for product roles
You describe the role once. They answer 12 written situations on their phone, in about 15 minutes. You get this report the second they finish — scored against your job, never against other candidates.
Free while we are in beta · ten credits · no card

Why this exists
By the final round everyone left is plausible. The portfolio shows what went well, always with a team behind it, and a one-hour interview mostly rewards whoever interviews best.
The take-home you stopped sending
Six hours of unpaid work filters by who has free time, not by judgment. Half your best candidates decline it, and the ones who accept produce artefacts you cannot compare.
The gut call you cannot defend
You know which one you liked. You cannot say why in a way that survives a conversation with your co-founder — or a complaint six months later.
The hire you got wrong once
Great portfolio, great interview, and it did not work. Nothing in the process was built to catch what actually went wrong.
Fifteen minutes of their time, and you get something you can put next to the interview: what this person reaches for first, and how close that is to what your job needs first.
Step one · about ninety seconds, once per role
Six capabilities, and they are the six that belong to this job — a UX researcher gets a different six. Mark the two or three the role genuinely needs and leave the rest as not relevant. What you rule out is what makes the comparison mean anything.

Step two · about 15 minutes of their time
No account, no install, nothing to prepare. Every option is something a sensible professional might do, so there is nothing to revise for and nothing to game.


Step three · the second they press submit
No queue, no grader, no waiting two days. The scoring is arithmetic with no language model anywhere in the path, so the report exists the moment the last answer lands.

A recommendation, not a percentile
It says what to do next and what to ask about. A number on its own moves the decision to you without helping you make it.
Three questions for the interview
Written from what this person actually answered. You walk into the call knowing where to push.
It tells you when not to trust it
Answers we cannot honestly read are withheld and the credit comes back. A number we do not believe is worth less to you than no number.
Running right here
Not an image. The block below is the real scoring engine, running on the server that served you this page.
Recommendation
Reaches first for spots when the brief is wrong. The distance here is not capability, it is how your team works.
Do they want the same things?
15 decisions behind it
You asked for it, they lead with it: spots when the brief is wrong, sweats states and edge cases.
Where each capability sits for them
| Dimension | Score out of 100 | What the role needs | |
|---|---|---|---|
| Spots when the brief is wrong | 71 | Essential | |
| Sweats states and edge cases | 63 | Essential | |
| Knows what is cheap to build | 54 | Not relevant | |
| Moves without being told | 20 | Not relevant | |
| Holds a position with a senior stakeholder | 13 | Not relevant | |
| Explains decisions outside the team | 21 | Not relevant |
You marked spots when the brief is wrong and sweats states and edge cases as essential, so those are what the number is about. No line in this report says whether this person is any good. It says how close they are to this job.
How it is scored
There is no person grading and no language model anywhere in the scoring path. The same twelve situations, the same arithmetic, every time — which is why the report exists the second a candidate presses submit, not in two days.
What changes between two companies is not the scoring, it is the weighting. You say which capabilities the role cannot do without; we do the rest identically for everyone.
And when a result cannot honestly be read — the same option position chosen over and over, or the whole thing finished faster than it can be read — we withhold it and give you the credit back. A number we do not believe is worth less to you than no number.
How it works
Marta R. answered once. Change who is hiring and watch the recommendation change — same person, same answers, different call. This is the real scoring engine, running here.
Marta R. is a constructed profile, not a real person — but the arithmetic below is the same code a paying customer gets.
Marta R.
Close to your job
Leads with all 3 of the things you named
What this role called essential
The mark is the middle of the scale.
Ask them to name something they wanted to build and dropped because of what it would cost.
Plainly
The second column is the useful one. A product that only says what it is leaves you to work out the limits, and working them out is where things get invented.
| Category | Is it? | Why |
|---|---|---|
| A situational judgment test | Yes | Twelve written situations. Every option is something a sensible professional might do, so what is measured is which one you reach for and which you rule out. |
| Scored against a specific job | Yes | The hiring company names up to three capabilities the role cannot do without, and those weight the score. The same candidate reads differently for two different jobs. |
| A skills test | No | There is nothing to be right about. No option is the correct one, which is why it cannot be revised for. |
| A personality test | No | Four style scales — pace, breadth, structure, autonomy — are shown and never scored. Neither end of any of them is better. |
| A cognitive or aptitude test | No | Nothing is timed and nothing has a solution. How long each situation takes is recorded only to detect answering without reading. |
| A coding test | No | The front-end and design engineer banks measure what someone decides to build and push back on, never whether the code compiles. |
| A take-home assignment | No | Nothing is produced and nothing is submitted for review. Fifteen minutes, identical for every candidate, so the results are comparable. |
| An AI interview | No | Nobody is recorded. No camera, no microphone, and no language model anywhere in the scoring path — scoring is deterministic arithmetic. |
| A proctored exam | No | No lockdown, no monitoring, no countdown. Candidates sit it in their own time and can pause. |
| A job board or a recruitment agency | No | No candidate profiles are published, no candidate is ranked publicly, and no candidate data is sold. |
What people ask us
Then you already know its three problems. It is not comparable — one candidate spends two hours and another spends the weekend, so you end up ranking availability. It loses your best people, because anyone with three live processes turns down a weekend of unpaid work. And it has no outside reference: you learn which one you preferred, not whether either is any good. This takes fifteen minutes of theirs and none of your design lead's.
It measures one thing, on purpose. Not craft, not tooling, not how they present — judgment. Twelve situations where every option is something a sensible professional might do, and what informs the profile is which action they reach for first and which they rule out. Ruling something out is the part a portfolio never shows you. If you want to see their craft, look at their work; that is what a portfolio is genuinely good for.
It buys them very little here. There is no essay to generate and no artefact to polish — every option is already written and defensible, so a model has nothing to improve on. What it cannot do is know which of four reasonable actions this particular person would take first. And we flag the patterns that show someone answered without reading: the same position chosen over and over, or the whole thing done faster than it can be read. Those results are withheld and your credit comes back.
No, and that is deliberate: an exam you can prepare for stops measuring anything. What is public is everything that matters for trusting it — the six capabilities that belong to each role, the four style scales, that the weights are yours rather than ours, and that no option is the correct one. You know what is measured. Nobody gets to rehearse how it is asked.
There is no person in the loop and no language model in the scoring path. The same twelve situations, the same arithmetic, every time — which is why the report appears the second a candidate finishes. What changes between two companies is not the scoring, it is the weighting: you say which capabilities this role cannot do without, and the same candidate reads differently for a different brief. That is the honest kind of subjective, and it is yours.
Fifteen minutes is not a test, it is a questionnaire. They are told before they start that there are no right answers, that the company in the situations is fictional so nobody is working for free, and that we measure how they decide rather than how they present. And it is once, not once per company: if another company asks for the same assessment in the next six months, they do not sit it again.
Then that hire matters more to you than to anyone, and getting it wrong costs a full salary and half a year. Your first candidate is free, with no card. Use it on the process you have open right now.
The beta
We are opening a small number of accounts to teams with a role open right now. What we need is not your money — it is what happened with the hire.
Ten credits
Ten candidates, on us. Credits never expire.
The full report
From the first candidate. There is no cut-down version.
No card
Nothing to cancel, because there is nothing to start.
All we ask back: mark in the panel what you decided about each candidate, and fifteen minutes on a call four weeks in. That is the data that will one day tell us whether this predicts anything — and until it does, we say so.
Take it yourself first if you like — the same 12 situations, about 15 minutes.
Related reading
9 min
Hiring a product designer well comes down to assessing decisions rather than screens: whether they reframe a badly posed brief, what they ask before designing, what they cut when resources disappear, how they respond when a constraint changes, and whether they can name the weakest part of their own proposal. Visual craft is necessary but it is the easiest thing to verify and the least predictive of seniority.
Read it10 min
Most design interview questions fail because a well-prepared candidate can answer them from their portfolio. Questions that discriminate share one property: they require information the candidate could only have if they were actually in the room — a decision that went wrong, a constraint they lost to, a thing they would cut.
Read it10 min
Self-report personality instruments measure what a candidate believes about themselves, or what they think you want to hear — which in a hiring context is the same thing. They predict job performance far more weakly than work samples, create legal exposure when used as a selection gate, and are visibly resented by experienced candidates.
Read it