Product23 Sep 20268 min read

The structured interview, from criteria to scoring sheet

Most hiring managers believe they can tell within ten minutes. The uncomfortable part is that they can: they can tell whether they enjoyed the conversation. Whether the person will do the job well is a different question, and a free-flowing chat is a poor instrument for answering it.

Why structure predicts better

Selection research has pointed the same way for decades. Interviews with a fixed question set, scored against criteria defined before anyone walked in, predict later job performance considerably better than open conversations. The reason is not mysterious: in an unstructured interview every candidate answers different questions, so at the end you are comparing impressions rather than evidence.

There is a second effect. An open conversation rewards candidates who resemble the interviewer: same background, same references, same way of talking. In the room that feels like rapport. On a scoring sheet it is noise.

Structure is not the same as rigidity. Four things are simply fixed before the first candidate arrives: what you measure, how much each measure counts, what you ask, and what a good answer sounds like.

Derive the criteria from the role, not from the CV

Criteria drawn from a pile of applications mostly describe the people who happened to apply. Start from the work. Write down what this person will spend their first year doing, then turn that into a short list.

  • List the five or six activities that take up most of the time in the role.
  • For each one, name the capability it needs in plain words: "explains a technical trade-off so a non-technical colleague can decide" rather than "communication skills".
  • Drop anything you cannot observe. Attention to detail as a personality trait is not assessable in an hour; catching the error in a sample report is.
  • Stop at four to six criteria. Beyond that the interview becomes a survey and every criterion gets less attention.
  • Check the list against the job advert. What is not in the advert either belongs there or nowhere.

One last pass: strike any criterion that is a proxy for background rather than capability. A named former employer, an unbroken career history or a particular university say more about someone's circumstances than about their work.

Weight the criteria before you read anything

Weights decide outcomes more quietly than scores do. Two candidates with the same total land in either order depending on whether analytical depth counts twice as much as stakeholder handling. Set the weights before the first application is opened.

A quick method that holds

  • Distribute 100 points across the whole set of criteria.
  • Nothing below 10 points. If a criterion is worth less than that, it is a side note.
  • Write one sentence per criterion explaining why it carries that weight.
  • Agree the weights in writing with whoever signs off the hire, before the first interview.

Once interviews begin, the weights stay put. Adjusting them after meeting people is how a panel talks itself into the candidate it already liked.

Two question types that do the work

Behavioural questions

These ask about something that already happened: "Tell me about a time you had to deliver with incomplete information. What did you do?" The value sits in the follow-ups. What exactly was your part, what did you decide, who disagreed, how did it end. Candidates arrive with a polished version of the story; two or three specific follow-ups get you to the real one.

Situational questions

These put a realistic problem from the role in front of the candidate: "A customer escalates two days before release and your most experienced developer is on holiday. Walk me through your first hour." They work well for people who have not held this exact role, because they test reasoning rather than access to past opportunities.

Ask every candidate the same core questions in the same order, and prepare the follow-ups too, so probing does not become improvisation that only some candidates face.

A scoring sheet with an anchored 1 to 5 scale

A bare 1 to 5 says very little, because your 4 is your colleague's 3. Anchor the levels: write down what a 1, a 3 and a 5 sound like in an answer. For "handles conflict with stakeholders", the anchors might read like this.

  • 1: describes the situation in general terms only, own role not visible, no concrete action named.
  • 3: describes a real case and their own part in it, but the outcome stays vague or was never measured.
  • 5: real case, own decisions named, trade-offs explained, states the result and what they would do differently today.

Score each criterion straight after the answer it belongs to, not at the end of the hour, and note the sentence from the answer that produced the score. A score with no quotation behind it is a feeling with a number in front of it.

Then multiply score by weight and add it up. The total orders the shortlist. It does not make the decision.

Keeping the halo effect and the first impression out

Two effects do most of the damage. The halo effect means a strong impression on one criterion bleeds into every other rating. The first-impression effect means the rest of the hour goes into collecting evidence for a verdict reached in the first two minutes. Neither disappears through experience or good intentions.

  • Score criterion by criterion across all candidates instead of filling in one whole form per person.
  • Have each interviewer score independently and submit before the debrief. Whoever speaks first turns three opinions into one.
  • Allow "culture fit" only if it is written as observable behaviour. Undefined, it usually means "like us".
  • Hold back the overall recommendation until every criterion has a score and a quotation.
  • Standardise the opening: same greeting, same order of questions, the same interviewer where possible.
  • Record disagreement instead of settling it in the corridor. A split panel is information.

None of this removes bias. It limits the places where bias can enter, and it leaves a trail you can find again later.

Documenting decisions so they hold up later

Six months on, nobody remembers why candidate B was preferred over candidate A. If the decision is questioned later, by a rejected applicant or an internal review, only a record written at the time helps. Keep it short and keep one per role.

  • The criteria and weights, dated, from before the first interview.
  • The question set every candidate received.
  • Per candidate: the score for each criterion, the quotation behind it, who scored and when.
  • The final decision, who made it, and the one or two criteria that actually decided it.
  • A retention rule for applicant data that covers these records too.

Write the reasons in the language of the criteria. "Scored 2 on handling conflict, named no concrete case when asked twice" is a reason. "Did not quite feel right for the team" is the kind of note that reads badly a year later.

How this runs in Persohap

Persohap follows the same order. Every applicant's CV is screened against criteria the recruiter defines and weights: the AI drafts a first set, the recruiter rewords and weights them before a single CV is read, and screening is blind to name, gender, age and origin. Only the candidates the recruiter picks are invited to an AI-led interview, which runs over one link with no account needed. The avatar opens with a line from the candidate's own CV and follows up on the answers actually given. Back come a transcript and a scorecard whose scores quote the transcript passage behind them, plus a recommendation of advance, hold or reject with a confidence value. A person makes the hiring decision; the system never decides.

See how structured hiring works in Persohap

Get started

See it working, not just described.

Every decision in these posts is visible in the product. We will walk you through whichever one you care about.