FirstHR

STAR Method of Interviewing: How Employers Score Answers

The STAR method of interviewing from the employer side: what a complete answer contains, how to probe for missing parts, and how to score consistently.

Nick Anisimov

Nick Anisimov

FirstHR Founder

Hiring
17 min

The STAR Method of Interviewing

Written for the person asking the questions, not the person answering them: what Situation, Task, Action and Result each prove, what a complete answer contains, how to write questions that produce one, how to probe the missing component when a candidate hides inside the word we, how to score the same answer the same way your colleague does, and the two things the method cannot tell you no matter how well you run it

I once hired somebody almost entirely on the strength of a single answer. I had asked about a time she rescued a project that was going badly, and I got nine minutes of vivid detail: the client, the deadline, the panic, the weekend, the recovery. I gave it a five out of five and I remember thinking the interview was over. Four months later it was clear she had been in the room for all of it and accountable for none of it.

The evidence was sitting in my own notes. She had said we more than thirty times and I twice. The answer had a Situation and it had a Result. What it did not have was a task that belonged to her or a single action she personally took, and I did not notice, because the story was good and I was enjoying it. That is the entire reason the STAR method exists on the employer side of the table.

Almost everything written about STAR is written for candidates: how to package your experience, how to sound structured, how to land the result. This is the other version. What each component actually proves, how to write questions that produce complete answers, how to probe the piece that is missing, how to score the same answer the same way a colleague scores it, and the two things the method cannot tell you no matter how carefully you run it. I build HR and onboarding software for companies with no HR department at FirstHR, and interview quality is the cheapest thing a small team can fix.

TL;DR
STAR stands for Situation, Task, Action and Result. As the interviewer you use it as a completeness check: one specific past event, a responsibility that belonged to the candidate, steps they personally took, and an outcome they can describe. Missing components are the signal. Probe for them with written follow-ups, then score against written anchors before you discuss anything.

What the STAR Method Is From the Interviewer’s Side

STAR is a completeness check you run on an answer, not a question format and not a coaching device. Situation, Task, Action and Result are the four things an account of past behavior has to contain before you are entitled to treat it as evidence of anything. Everything else in this article follows from that one idea.

Definition
STAR method
A framework for evaluating answers to behavioral interview questions by checking whether the answer contains four components: the Situation the candidate was in, the Task or responsibility that was theirs, the Actions they personally took, and the Result those actions produced. Interviewers use it in two ways. First as a scoring lens, marking which components an answer contains. Second as a probing structure, because the missing component tells you exactly which follow-up question to ask.

The reframe that matters for a small employer is this. Most people treat STAR as a way to help good candidates express themselves. It is more useful as a way to stop yourself being persuaded. An answer that sounds impressive and contains no personal Task is the single most common way an interviewer talks themselves into a hire that fails in month four.

It is also worth remembering what an interview legally is. Under the Uniform Guidelines on Employee Selection Procedures, a selection procedure means any measure or procedure used as a basis for an employment decision, and the definition explicitly reaches the full range of assessment techniques through to informal or casual interviews (29 CFR 1607.16). Your interview is a test. STAR is one of the few practical ways a company without an HR department can make that test consistent.

Situation, Task, Action and Result: What Each One Is Evidence Of

Each component answers a different question, and each one fails in a characteristic way. Situation establishes that a real event happened. Task establishes ownership. Action establishes contribution. Result establishes effect. Only the combination is evidence.

SSituation: the specific event, with a time and a place
One identifiable moment at one employer, not a summary of how the candidate generally operates.Missing when the answer is written in the present tense: I always make sure the team knows where things stand. That is a philosophy, not an event, and there is nothing in it you can verify.
TTask: what the candidate personally owned
The obligation that sat on this person. What were they on the hook for, and what happened if they got it wrong?The most commonly missing component and the most expensive one to miss. Without it you cannot tell a project owner from a person who watched the project happen from two desks away.
AAction: the steps they took, in the first person singular
Sequence, choices, and the reasoning behind them. Who they spoke to, what they decided, what they tried that did not work.Missing when actions belong to a group. A candidate who cannot name three things they personally did is describing something they were near rather than something they did.
RResult: what changed, and how they know
An outcome with a number, a date, a decision, or a consequence. A failed outcome the candidate can explain is worth more than a vague success.Missing when the answer ends on effort rather than effect: and after that we were all much more aligned. Ask what became true afterwards that was not true before.
Candidates rarely deliver the four components in order, and they do not have to. Your job is presence, not sequence. Tick each letter as you hear it and treat the blanks as the interview’s real agenda.

Two clarifications save a lot of confusion in practice. Candidates almost never deliver these in order, and they do not need to. A perfectly good answer might open on the result and work backwards. Score presence, not sequence.

And a Result does not have to be a success. A candidate who describes a project that failed, explains what the failure cost, and says what they would do differently has given you a complete and highly informative answer. Some of the strongest answers I have scored were about things that went badly. What you are grading is whether they can connect their actions to an effect in the world, not whether the effect flattered them.

What a Complete Answer Looks Like Next to a Weak One

A complete answer names one event, states what the candidate owned, describes at least three specific things they did, and reports a verifiable outcome. A weak answer usually has a Situation and a Result and nothing solid in between. Here are both, for the same question.

The question: tell me about a time you had to fix a problem that somebody else on your team created.

ComponentThe weak answerThe complete answer
SituationWe had a lot of issues with our invoicing last year and things were pretty chaotic on the finance sideIn March at my last job, a colleague sent forty-one invoices to the wrong billing contacts after a CRM import went wrong. Two clients called our CEO the same afternoon
TaskSo we all pitched in to sort it out because that is how our team worksI owned the billing inbox, so it was my call how we corrected it and my responsibility to keep the two escalated clients from churning
ActionWe reached out to the clients, we fixed the data, and we made sure it did not happen againI pulled the export and identified the forty-one records in about an hour. I called the two escalated clients myself before end of day rather than emailing. I credited one account for the confusion after checking with my manager. Then I wrote a two-step check into our import process and walked the colleague through it, because he was going to run the next import
ResultIt got resolved and everyone was happy in the endBoth clients stayed, one renewed at a higher tier that June. We ran nine more imports before I left with no repeat errors. It took me about a day and a half in total
What you can scoreNothing. There is no ownership, no personal action, and no outcome you could check with a referenceOwnership, judgment about escalation, a decision made with a manager, a process fix, and a result with dates and numbers attached

Notice what makes the second column strong, because it is not eloquence. It is countable specifics: forty-one records, two escalated clients, one credit, nine subsequent imports. Specifics are the part of an answer that a reference check can confirm or contradict, which is why the two stages belong together.

Notice also what makes the first column weak. It is not short. Weak answers are often longer than strong ones, because generality expands to fill whatever time you give it. Length is the least reliable quality signal in an interview.

How to Write Questions That Produce STAR-Shaped Answers

A question produces a STAR-shaped answer when it asks for one specific past event and names a single behavior you care about. Questions in the present tense produce philosophy. Questions in the conditional produce speculation. Both are worth asking sometimes, but neither can be scored with STAR.

Question that will not produce STARWhy it failsRewrite
How do you handle conflict with a coworker?Present tense and general. Invites a self-description the candidate has rehearsed and you cannot verifyTell me about a specific disagreement you had with a coworker in the last two years. What was it about and how did it end?
Are you a good communicator?Closed and leading. There is one answer and everyone gives itDescribe a time you had to explain something technical to somebody who did not have the background. What did you do to make it land?
What would you do if a customer threatened to cancel?Hypothetical. There is no real situation and no real result, so there is nothing to checkTell me about a time a customer told you they were leaving. Walk me through what you did and what happened
Tell me about a time you showed leadership, teamwork and initiativeThree competencies in one question. Whatever the candidate answers, you cannot say which one you scoredTell me about a time you got a group to do something they were reluctant to do. What was your role in it?
Tell me about your greatest weaknessNot behavioral at all. Measures interview preparation and nothing elseTell me about a piece of feedback that was hard to hear. What did you do in the following month?
You seem like a self-starter, can you give me an example?Leading. You have told the candidate the answer before they speakTell me about something you started at work that nobody asked you to start. What happened to it?

Four rules cover almost everything. Past tense. One event. One competency, drawn from the job description rather than invented in the room. And no adjectives that tell the candidate which answer you are hoping for.

Ask the same questions in the same order of every candidate for the role. That is what makes the scores comparable, and it is also what keeps you away from the topics that create legal exposure, which are set out in the guide to illegal interview questions.

One structural point about mixing question types. Behavioral questions ask what a candidate did; situational questions ask what they would do. Both belong in a well-designed interview and each has its own scoring approach, which is why situational interview questions come with their own rubric rather than being forced through STAR.

Still Using Spreadsheets for Onboarding?
Automate documents, training assignments, task management, and track onboarding progress in real time.
See How It Works

How to Probe When a Component Is Missing

Probe the specific missing component, using a follow-up you wrote before the interview, and cap yourself at two probes per question. The probe is not a courtesy to the candidate. It is how you find out whether the gap in the answer is a communication habit or an absence of experience.

No Situation: the answer is a general policyAsk for one instance. Give me the most recent time that actually happened. Where were you working, roughly when was it, and who else was involved? If the candidate cannot land on a single event after one probe, score the answer on what you have rather than helping them invent one.
No Task: you cannot tell what they ownedAsk about accountability, not activity. What part of this were you responsible for delivering? Who would have been asked about it if it had gone wrong? A candidate with real ownership answers this in one sentence. A candidate without it changes the subject back to the team.
No Action: the story stays at altitudeForce the sequence. Walk me through what you did first, then what happened next. What did you decide that somebody else might have decided differently? Ask for the thing they tried that did not work. Invented stories almost never contain a failed first attempt.
No Result: the story ends on effortAsk what changed. How did it turn out, and how did you find out? What would the number have been if you had not done anything? If the result is unknown even to the candidate, that is itself informative about how their previous employer measured work.
Everything is we and never IName it neutrally and split the group. It sounds like several people were involved. Take me through your specific piece of it. If the candidate keeps returning to the group after two attempts, stop probing and score the answer as unevidenced rather than as bad.
The answer is long and still emptyInterrupt politely and reset. Let me pull you back to one thing. Length is not evidence, and a nine-minute answer with no Task and no Action scores lower than a ninety-second answer that has all four. Give every candidate the same amount of runway.
Write your probes down before the interview and cap them at two per question. Improvised follow-ups are where consistency quietly dies, because the candidate you like gets four helpful prompts and the candidate you do not gets none.

The two-probe cap does real work. Without it, interviewers probe hardest for the candidates they already like, effectively coaching them into a better score, and give up quickly on the ones they do not. That asymmetry is invisible from inside the room and obvious in the scores afterwards.

Take notes in the candidate’s language rather than your conclusions. Writing down good energy, seems organized gives your debrief nothing to work with. Writing down the phrase they used, the number they quoted and the pronoun they chose lets a colleague who was not in the room re-score the answer from your notes. That is the test of a usable interview note.

One more thing about the probe itself: ask for the part that went wrong. What did you try first that did not work is the highest-yield follow-up I know. Real experience contains dead ends. Borrowed or embellished stories are usually clean, linear and slightly too tidy, which is a pattern worth reading alongside the other interview red flags.

The We Problem: Separating the Candidate From the Team

When a candidate says we throughout an answer and never I, you have no Task and no Action, which means you have nothing to score. Ask once for their specific piece, ask a second time if needed, and then treat the answer as unevidenced rather than as a character finding.

I want to be careful here, because this is where interviewers over-read. Collective language is not proof of dishonesty. Some organizations actively train people to attribute work to the team. Some candidates come from cultures where claiming individual credit is impolite. Nervous people default to the group. None of that means the person did nothing.

The neutral split works better than a challenge. It sounds like several people were involved, so take me through your specific piece of it, invites the candidate to answer rather than defend themselves. If they still stay collective, try the accountability angle instead of the activity angle: if this had gone badly, who would your manager have called? Ownership questions are harder to deflect than contribution questions.

Read the Pattern, Not the Answer
Track pronouns across the whole interview rather than reacting to one answer. A candidate who produces a clear personal Task in four questions and goes collective on the fifth is probably describing a project where they genuinely were a supporting player, which is honest. A candidate who never lands a personal action across five questions has given you a consistent finding. Score each answer on its own evidence, then look at the shape of the set at the debrief. The pattern is the signal; a single we is noise.

There is a mirror-image failure worth naming. Some candidates say I about work that was obviously collective, taking sole credit for a launch that took eleven people. Probe that the same way, from the other direction: who else was involved, and what did they do? An answer that cannot describe anybody else’s contribution is as incomplete as one that cannot describe the candidate’s own.

Scoring STAR Answers on an Anchored Rating Scale

Consistent scoring comes from written anchors, not from experience or good intentions. For each question, write down what a low, middle and high answer actually contains before anybody interviews, then have every interviewer score independently and in writing before the debrief starts.

Anchors describe evidence, never impressions. The difference between a scale that works and one that does not is whether the descriptions could be applied by somebody who has never met the candidate and is reading only your notes.

RatingWhat the answer containedHow to recognize it
1: No evidenceNo specific situation, or a situation with no personal task and no personal actionGeneral statements about how the candidate usually works. Survives two probes without producing a single event
2: Weak evidenceA real situation, but ownership is unclear and actions are collective or genericThe candidate names an event but the work belongs to a group. Result is vague or absent
3: Adequate evidenceAll four components present at a basic level. One event, a personal task, at least two actions, an outcomeComplete but thin. You could describe what they did in one sentence and could not say why they chose it
4: Strong evidenceComplete, specific, and includes reasoning. Actions have a sequence and at least one judgment callThe candidate explains a decision, mentions an alternative they rejected, and gives a result with a number or a date
5: Exceptional evidenceEverything in 4, plus difficulty, self-correction and transferThey describe what went wrong first, what they changed, and what they carried into the next situation. The complexity is above the level of the role

Five points is a reasonable default. The Office of Personnel Management guide to structured interviews recommends labeling at least three proficiency levels and using five to seven where you can define them properly (OPM, Structured Interviews: A Practical Guide, September 2008). More points only help if you can genuinely tell them apart in writing. If you cannot describe the difference between a 6 and a 7, you do not have a seven-point scale, you have a five-point scale with noise.

1
Write anchors per question, not per competency
A generic five-point scale for communication is useless. The anchors have to describe what a strong answer to this particular question contains, which takes about ten minutes per question and is written once per role.
2
Score in the room, before any discussion
Write the number on the form before you stand up. Scores announced out loud first pull everybody else toward them, and that anchoring effect is stronger than most interviewers believe.
3
Score components, then the answer
Tick which of S, T, A and R were present, then assign the rating. The component ticks are what let two interviewers work out why they disagreed.
4
Never adjust one score to fit your overall impression
If a candidate scored a 2 on delegation and a 5 on customer recovery, that is the finding. Smoothing the low score because you liked them destroys the reason for scoring at all.
5
Compare question by question at the debrief
Discuss only the questions where scores differ by two or more points, and settle them by re-reading notes rather than restating opinions.
6
Calibrate the scale on real answers once a quarter
Take one recorded or written-up answer, have everyone score it independently, and compare. Half an hour of this does more for consistency than any amount of interview training.
Why Structure Beats Judgment, in Numbers
The research on interview structure is unusually consistent. McDaniel, Whetzel, Schmidt and Maurer’s meta-analysis in the Journal of Applied Psychology (1994), covering 245 validity coefficients across 86,311 individuals, found mean corrected validity of 0.44 for structured interviews against 0.33 for unstructured ones. Conway, Jako and Goodman, in the same journal (1995), found average interrater reliability of 0.59 for highly structured individual interviews against 0.37 for unstructured, with standardization of response evaluation as a key moderator. Sackett, Zhang, Berry and Lievens (Journal of Applied Psychology, 2022) put structured interviews at 0.42 in their reanalysis, ahead of cognitive ability tests at 0.31.

Two of those three findings are about how answers are evaluated rather than how questions are asked. That is the part small employers skip. Writing standard questions takes an afternoon and feels productive; writing anchors is duller and produces most of the gain.

Companies Using FirstHR Onboard 3x Faster
Join hundreds of small businesses who transformed their new hire experience.
See It in Action

The Failure Almost Nobody Catches: Rewarding Storytelling Over Evidence

The most common way STAR goes wrong is that interviewers score the quality of the telling instead of the quality of the evidence. A candidate with a rehearsed, well-paced, emotionally satisfying story gets a five for a project they observed, and a candidate who is nervous and disorganized gets a two for work they personally led.

This happens because narrative competence and job competence feel like the same thing in the moment. Both produce the sensation of watching somebody who knows what they are doing. They separate cleanly on paper and almost never in the room, which is exactly why you write the score against anchors instead of trusting the feeling you leave with.

Three Symptoms That You Are Scoring the Story
Check your last five scorecards for these. First, your highest-scoring answers are also your longest ones, which suggests you are rewarding airtime. Second, your notes contain adjectives about the person (confident, articulate, impressive) rather than quotes about the work. Third, you cannot state, from your notes alone, what the candidate personally decided in their best answer. If a hiring decision rests on an answer you cannot reconstruct from evidence, you did not run a STAR interview. You watched a presentation and gave it a grade.

There is a second version of this failure that runs the opposite way and costs you good people. Because STAR rewards a specific verbal shape, candidates who have been coached on it perform better than candidates who have not, independent of ability. Interview coaching is unevenly distributed. Telling every candidate the format at the start of the interview is the cheapest correction available, and it takes about twenty seconds.

The underlying problem is that most companies never find out whether their interviews work. Peter Cappelli made this the center of his argument in Harvard Business Review (May to June 2019), that employers invest heavily in hiring methods while rarely measuring whether those methods produce good employees (Your Approach to Hiring Is All Wrong). The small-business version of the fix is modest: keep the scorecards, and once a year compare interview scores against how those hires actually turned out.

What the STAR Method Cannot Assess

STAR measures reported past behavior, so it cannot demonstrate a skill and it cannot predict judgment in a situation the candidate has never been in. Those two gaps are structural. No amount of probing closes them, and every well-run process covers them with something other than an interview.

4
components an answer must contain before it counts as evidence
0.44
mean corrected validity of structured interviews, McDaniel et al. meta-analysis (1994)
0.59
interrater reliability of highly structured interviews vs 0.37 unstructured (Conway et al., 1995)
5 to 7
rating levels recommended by the OPM structured interview guide (2008)

The first gap is skill demonstration. A beautifully evidenced account of writing a marketing plan is a report about writing, not a piece of writing. If the job turns on a craft, watch the craft. A short paid exercise, a work sample, or a structured trial task tells you in ninety minutes what six behavioral questions cannot.

The second gap is judgment under novel conditions. STAR asks what somebody did in situations they have already met. It is close to silent on what they will do in situations they have not, which is precisely the question you are asking when you promote an individual contributor into management, hire the first person in a function, or bring somebody into a company far smaller or larger than the one they came from. Hypothetical and case-style questions exist to cover that gap, and they are scored on reasoning quality rather than on STAR.

Three smaller limits are worth knowing. STAR is weak with early-career candidates, who have thin professional histories and will either strain to fit school projects into the frame or produce nothing. It rewards recall and verbal fluency, which are only sometimes part of the job. And it is easy to over-index on recent experience, since the most vivid stories are usually the newest ones, not the most relevant.

Where STAR Fits in a Structured Interview Process

STAR is one component of a structured interview, not a substitute for one. Structure is the whole system: the same job-relevant questions in the same order for every candidate, written scoring anchors, independent scores, and a debrief that compares evidence. STAR is the specific technique you apply to the behavioral questions inside that system.

If you have not built the surrounding process yet, build that first, because STAR applied inside an unstructured conversation adds very little.

A practical sequence for a small team looks like this. A short screening call on non-negotiables. One structured interview with four to six behavioral questions scored with STAR. A work sample or trial task where the role has a craft. One conversation with the person the candidate would report to, still on written questions. Then references.

Train whoever else interviews. Two hours is enough: explain the four components, run one practice answer that everybody scores independently, compare the numbers, and argue about the gaps. Most interviewer disagreement comes from different definitions of a four, not from different readings of the candidate, and that is fixable in an afternoon.

Keep the paperwork. The EEOC advises employers to ensure that selection procedures are properly validated for the positions and purposes for which they are used, and to be able to show that a procedure is job-related and consistent with business necessity (EEOC guidance on employment tests and selection procedures). Completed scorecards with contemporaneous notes are the most persuasive record a small business can produce, and they are also what makes useful interview feedback possible.

The last piece is the loop nobody closes. Six months after each hire, pull their interview scorecard and ask whether the scores predicted anything. You will find that one or two of your questions consistently separate strong hires from weak ones and the rest do not. Keep those questions, rewrite the others, and your interview gets better every year instead of staying exactly as good as it was the day you wrote it.

Key Takeaways
STAR is a completeness check the interviewer runs on an answer, not a storytelling formula for candidates. Situation, Task, Action and Result are the four things an answer must contain before it counts as evidence.
Task is the component most often missing and the most expensive to miss, because without it you cannot separate a person who owned the work from a person who was present for it.
A Result does not have to be a success. A failure the candidate can explain, with a cost and a lesson, is a complete and highly informative answer.
Questions produce STAR answers only when they ask for one specific past event and one competency. Present-tense questions produce philosophy and hypothetical questions produce speculation.
Write two probes per question in advance and cap yourself there. Improvised follow-ups mean the candidates you like get coached into higher scores.
When a candidate says we and never I, split the group once, ask the accountability question once, then score the answer as unevidenced rather than concluding anything about their character.
Consistency comes from written anchors that describe evidence, not impressions. Score independently and in writing before anybody says a number out loud.
The most common failure is rewarding storytelling over evidence. If your highest scores are also your longest answers, you are grading a presentation.
STAR cannot demonstrate a skill and cannot predict judgment in a situation the candidate has never faced. Cover those gaps with work samples and scenario questions.
Research is consistent on the payoff: McDaniel et al. (1994) found 0.44 validity for structured interviews against 0.33 unstructured, and Conway et al. (1995) found interrater reliability of 0.59 against 0.37.

Frequently Asked Questions

What is the STAR method of interviewing?

STAR stands for Situation, Task, Action and Result, and it is a framework an interviewer uses to judge whether an answer to a behavioral question contains real evidence. A complete answer names one specific past event, states what the candidate personally owned in it, describes the steps they took in the first person, and reports what changed as a result. The value of the framework for an employer is not that it makes candidates sound polished. It is that the missing pieces are diagnostic. An answer with a vivid Situation and no Task tells you the candidate may have been present rather than responsible. An answer that ends on effort rather than outcome tells you nobody measured their work, or that they never asked. You score what the answer contains and you probe for what it does not.

Is the STAR method only useful for behavioral questions?

Mostly yes, and that is worth being precise about. STAR evaluates accounts of past behavior, so it fits questions that begin with tell me about a time or describe a situation where. It does not fit hypothetical questions, which ask what a candidate would do in a scenario you invent, because there is no real Situation and no real Result to report. Hypothetical questions are still worth asking and have their own scoring approach based on the quality of the reasoning. It also does not fit knowledge checks, work samples, or any exercise where the candidate demonstrates the skill in front of you. Use STAR as one lane in the interview rather than the whole road, and be clear with your interviewers about which questions are being scored which way.

How do you score a STAR answer consistently across interviewers?

Write the rating anchors before anybody interviews, and write them about evidence rather than impressions. For each question, state in one or two sentences what a low, middle and high answer actually contains: which components are present, how specific the actions are, and whether the result is verifiable. The Office of Personnel Management structured interview guide recommends labeling at least three proficiency levels and using five to seven levels where you can. Then have every interviewer score independently and in writing before the debrief. Conway, Jako and Goodman’s meta-analysis in the Journal of Applied Psychology (1995) found average interrater reliability of 0.59 for highly structured individual interviews against 0.37 for unstructured ones, and most of that gap comes from standardizing how responses are evaluated rather than from the questions themselves.

What should you do when a candidate says we and never I?

Name it neutrally once, split the group, and then stop. Something like: it sounds like several people were involved, so take me through your specific piece of it. Most candidates are not hiding anything. Plenty of workplaces train people out of taking individual credit, and some cultures treat first person singular as boasting. That is why you probe rather than conclude. If two attempts still produce collective language with no personal task and no personal action, the correct outcome is to score the answer as unevidenced, not to score the candidate as dishonest. Move on to the next question, which gives them a fresh chance. A candidate who cannot produce a personal contribution across several questions is telling you something consistent, and that pattern is the finding, not any single answer.

Should you tell candidates you are using the STAR method?

Yes, tell them at the start of the interview. Say that you will ask about specific past situations, that you are interested in what they personally did, and that you will take notes and may interrupt to ask for detail. This costs you nothing and removes a source of noise you do not want to measure. Candidates who have been coached on STAR already know the format, so staying silent only advantages the ones with better preparation, which is rarely the thing you are hiring for. Telling everyone equalizes that. It also makes probing feel collaborative rather than adversarial, which produces longer and more honest answers. The one thing not to share in advance is the rating anchors, since a candidate who knows exactly what a five contains can aim at it without having lived it.

What can the STAR method not tell you about a candidate?

Two things above all. It cannot tell you whether somebody can actually do the work, because a well-told account of past performance is a report about a skill rather than a demonstration of one. If the job involves writing, building, selling or fixing something, you need a work sample or a paid trial task to see the skill itself. And it cannot tell you how somebody will handle a situation they have never faced, which matters most for first-time managers, first-time founders’ hires, and anyone stepping up a level. STAR is backward looking by construction. It is also weak with early-career candidates who have little professional history to draw on, where scenario questions and a structured trial task give you far more than a hunt for past examples that do not exist yet.

How many STAR questions should one interview include?

Plan for four to six in a sixty-minute interview and resist the urge to add more. A complete answer with two probes takes six to eight minutes once you include the candidate’s thinking time, so six questions fills forty-eight minutes before you have said hello or left room for their questions. Interviewers who schedule ten behavioral questions end up rushing the probes, and the probes are where the method earns its keep. Choose the competencies that the job actually turns on, cover each one with a single question, and split the rest across other stages if you have them. It is better to have four questions you scored properly than ten you skimmed, because an unprobed answer contributes noise to the decision rather than signal.

Ready to transform your onboarding?

7-day free trial No credit card required
Start Your Free Trial