What 30 years of research says about structured interviews

The case for structured interviewing has been settled in industrial and organizational psychology for more than three decades. Newer research has not overturned the conclusion. It has refined it. The wording, format, and integration of structured interviews matters more than the field once thought — and small design choices change outcomes.

This is a short synthesis of where the evidence stands today and what it implies for executive hiring.

The core consensus

Structured interviews are more valid predictors of job performance than most other selection tools (Huffcutt & Murphy, 2023 [EN]). They produce more reliable and consistent evaluations than unstructured interviews (Segal et al., 2006; Stevens, 2012 [EN]).

The effectiveness comes from standardization. Same questions. Same rating scales. Same evaluation criteria. Standardization reduces bias and removes irrelevant influences from the decision. The advantage has been observed consistently across decades, though the exact psychological mechanisms remain incompletely explained (Dipboye, 2017 [EN]).

The headline finding is durable. The why is still being studied.

Fairness and reduced bias

Structured interviews predict performance well and show lower adverse impact on racial groups than other common predictors, including cognitive ability tests and assessment centers (Huffcutt & Murphy, 2023 [EN]). Increasing the level of structure minimizes the influence of factors unrelated to job performance (Stevens, 2012 [EN]).

For organizations operating under the EU AI Act, this finding is operationally significant. The Act requires high-risk hiring systems to demonstrate that their outputs are not producing discriminatory outcomes. Structured assessment, properly documented, supplies the evidence base for that demonstration.

Question design — an emerging frontier

A newer and less-explored area concerns how questions are written. Small differences in wording can significantly affect what the interview actually measures (Huffcutt & Murphy, 2023 [EN]).

A key distinction is between typical and maximal performance. Typical performance is what someone does on an ordinary day. Maximal performance is what they can do at their best. The two are weakly correlated. If the question does not make clear which one is being asked about, the answer measures neither cleanly. Unclear wording introduces measurement error and variability in validity.

Recent research suggests that priming participants — asking them to describe their most effective behavior in a given situation, rather than a generic example — improves consistency when assessing maximal performance.

The implication is direct. Even structured interviews can vary in quality depending on how precisely questions are written. Structure is necessary. It is not sufficient.

Format matters: past behavior vs situational

Not all structured interviews are equally effective.

Past Behavior Interviews (PBI) — which ask the participant how they actually behaved in past situations — show significant predictive validity for job performance, with reported correlations around r = .32 (Krajewski et al., 2006 [EN]). They are also more strongly related to relevant constructs like cognitive ability and personality.

Situational Interviews (SI) — which ask how the participant would behave in a hypothetical future scenario — show weaker or non-significant predictive validity in the same study.

The implication: asking about actual past behavior tends to be more predictive than asking about hypothetical future scenarios.

Kettrion’s question architecture defaults to past-behavior framing for this reason. Hypothetical framing remains available where it is more appropriate to the competency being assessed, but it is the exception, not the default.

Beyond hiring: clinical and high-stakes selection

The evidence base is not confined to hiring.

In clinical psychology, structured interviews improve diagnostic reliability and consistency. They produce more valid and meaningful diagnoses (Segal et al., 2006 [EN]). They support standardized, evidence-based practice in settings where treatment decisions carry significant consequences.

In medicine, structured competency-based interviews are considered essential for robust selection decisions. Effectiveness depends on three things: a clear competency focus, standardized rating systems, and integration with other assessment tools such as assessment centers (Jackson et al., 2007 [EN]).

The pattern is consistent. Where the cost of a wrong decision is high, structured assessment is the standard. Executive hiring carries one of the highest costs per failed decision in any people-related domain. The standard has not yet been applied at scale.

What the research has not yet resolved

Despite strong evidence favoring structured interviews, the field acknowledges that the mechanism — why they work so well — is not fully understood. Existing research often lacks integration of the social and cognitive processes that play out inside the interview itself (Dipboye, 2017 [EN]). Question design effects, including wording, remain an understudied but consequential frontier.

This matters for product design. The honest position is that structure is the strongest predictor available, and the marginal gains from better question design, format choice, and rater calibration are still being mapped.

What this implies for Kettrion

Kettrion is built on the consensus, not the frontier. The methodology defaults to past-behavior framing, anchored rating scales, structured rubrics, and observation-level data capture. Where the research is unsettled — wording effects, the typical-versus-maximal distinction, the mechanism question — Kettrion is built to surface the data that will help resolve it.

The corpus is an asset for the buyer first. Over time, it is also the substrate for refining the methodology itself. Every mandate makes the next one sharper.

The science has existed for 30 years. The software has not.

References

← All insights