Skip to content

How to score a design exercise

By Sergio Gualda · 9 min read · Updated

Score a design exercise against written anchors, not impressions: define three to five criteria, describe what each score level looks like in observable behaviour, fix the weights before you read anything, and require a quote from the candidate's work to justify every score. Then calibrate on people whose ability you already know.

Write the rubric before you send the exercise

If you cannot describe what a strong answer looks like before anyone has answered, the exercise is not ready. Writing the rubric first is also the cheapest way to discover that your prompt does not actually elicit the thing you care about.

A useful test: if two of your criteria would always be scored the same way by the same answer, you have one criterion, not two.

Anchors describe behaviour, not quality

The most common failure is a rubric whose levels are “excellent / good / adequate / poor”. That is not a rubric, it is a scale with adjectives, and two reviewers will apply it differently within minutes.

An anchor should describe something you can observe. Compare “excellent problem framing” with “rejects the brief as posed, reframes it in terms of the stated outcome, and cites at least two pieces of data from the material”. Only the second can be applied consistently.

  • 3 — Rejects the brief as posed and supports the reframe with two or more specific data points.
  • 2 — Notices the brief is wrong but argues from intuition, or uses a single data point.
  • 1 — Executes the brief and adds an improvement of their own.
  • 0 — Executes the brief literally. Never questions the premise.

Fix the weights before you read anything

Weights decide what the exercise is actually measuring, and deciding them after reading answers means the strongest candidate defines the rubric.

Weight towards what fails hires. In most product roles, misjudging the problem causes more damage than executing imperfectly, so problem framing should usually carry more weight than craft.

Require evidence for every score

Make the reviewer paste a quote from the candidate's answer next to each score. It takes a few extra minutes and changes the review completely.

It forces scores to be grounded in what was written rather than in an impression formed in the first two minutes. It makes disagreement between reviewers productive, because you are arguing about a specific sentence. And it produces feedback you can give the candidate.

Calibrate before you use it on strangers

Take three people you know are strong and three you know are weak — people you have actually worked with, not reputations. Have them do the exercise. Have someone anonymise the answers, then score them blind.

The rubric passes if the three strong ones score above the three weak ones with no overlap, and by a clear margin. If a weak one outranks a strong one, the problem is your rubric, not the candidate. Fix it and repeat.

Almost nobody does this, which is why almost every internal rubric quietly measures presentation.

Three failure modes to watch for

These show up in nearly every rubric written in a hurry.

  • Scoring effort. Length and polish are the easiest things to see and the least predictive. Say explicitly that they do not count.
  • Halo from the first section. Read all candidates' answers to section one, then all of section two, rather than one candidate end to end.
  • Criteria that measure the same thing twice, which silently doubles that criterion's weight.

Frequently asked

How many criteria should a rubric have?
Three to five. Beyond that, reviewers stop distinguishing between them and start assigning a general impression across all of them, which defeats the purpose.
Should candidates see the rubric?
Yes. There is no correct answer to leak, and knowing what is measured does not let anyone fake it — they still have to do it. What it does remove is the sense of ambush, which is the main reason strong candidates decline exercises.