The Instrument on Day Twelve
A review of the selection-validity evidence, and what it showed me about where an instrument belongs in a hiring process
I care about building systems that are fair. Not fair in the sense of feeling even-handed, which costs nothing and proves less, but fair in the structural sense: the same process for every candidate, recorded, repeatable and defensible afterwards. That conviction is most of the reason I do this job rather than a different one.
This week I set out to answer a question that ought to be straightforward for a company in our line of work. Are psychometric instruments actually valuable in hiring, and what does ignoring them cost?
I had three reasons. The people who use our instruments deserve to know what the evidence behind them actually says, including where it is weaker than the marketing suggests. Work like this feeds our own promotional material, and I would rather that material rest on numbers I have checked than on numbers that flatter us. And, less respectably but no less truthfully, this is the interesting part of the job.
The review did what I asked of it. It also turned the question around on me.
The line I wrote myself
I have spent recent weeks drafting an internal handbook for our commercial function. Part of it covers how new people join the business development team and find their feet.
One element of that induction I wrote with some satisfaction. Take the TEIQue Full Form personally and receive your own report, because nobody should sell an instrument they have not experienced, or one they do not believe works. We sit down with the report and work through why the scores came out as they did. The point is to make sure the instruments our business development executives sell are valuable in their own eyes. If they see no value in it, there is no sense in asking them to sell something they think is useless.
I still think that is right.
But reading the sequence in order, against the evidence I had just commissioned, something obvious surfaced. That step sits at roughly day twelve. Day twelve is after the conversations. After the decision. After the engagement is agreed. The instrument was not in the selection process at all. It was in the induction pack.
What follows is the review itself. I have put my own position at the end, after the evidence rather than in front of it, because that is the order in which the reasoning should run: evidence first, position after. The design was still on paper when the evidence arrived, which is when a review is worth most. Nothing had been executed, and the sequence could simply be put right.
Part one: what the evidence says about validity
The numbers changed, and changed downwards
Anyone citing selection-validity figures needs to know that the standard reference points were revised recently, and revised downwards.
For roughly two decades the anchor was Schmidt and Hunter (1998), a synthesis of eighty-five years of research published in Psychological Bulletin. Their operational validities made general mental ability look close to decisive.
| Method | Operational validity |
|---|---|
| Work-sample tests | ≈ .54 |
| General mental ability | ≈ .51 |
| Structured interviews | ≈ .51 |
| Job-knowledge tests | ≈ .48 |
| Integrity tests | ≈ .41 |
| Unstructured interviews | ≈ .38 |
| Conscientiousness | ≈ .31 |
| Years of job experience | ≈ .18 |
| Years of education | ≈ .10 |
They reported that general mental ability combined with a work sample, an integrity test or a structured interview produced composite validities of roughly .63 to .65. Those figures still appear in a great deal of vendor material across the category.
Sackett, Zhang, Berry and Lievens (2022) identified a systematic error in how they were produced.
Observed validity in any real sample is attenuated by two things. Criterion unreliability, because supervisor ratings are themselves imperfect measures of performance. And range restriction, because validation samples usually consist of people already selected into the job, who therefore vary less than applicants do. Correcting for both is legitimate in principle.
The problem Sackett and colleagues identified is that standard practice took large range-restriction corrections derived from predictive studies, where restriction is often substantial, and applied them to concurrent studies, where they demonstrate it is usually small. Because roughly eighty per cent of the studies in these meta-analyses are concurrent, the overcorrection propagated across the entire literature. Their proposed remedy is a principle of conservative estimation: absent a credible estimate of the ratio between applicant and incumbent standard deviations, do not correct at all, on the grounds that underestimating validity is a smaller error than overclaiming it.
The revised estimates reorder the field:
| Method | Schmidt & Hunter (1998) | Sackett et al. (2022) | Lower 80% credibility value |
|---|---|---|---|
| Structured interviews | .51 | .42 | .18 |
| Job-knowledge tests | .48 | .40 | .23 |
| Empirically keyed biodata | .35 | .38 | .26 |
| Work-sample tests | .54 | .33 | .21 |
| General mental ability | .51 | .31 | .13 |
| Integrity tests | .41 | .31 | .05 |
| Personality-based EI | not reported | .30 | .08 |
| Assessment centres | .37 | .29 | .17 |
| Contextualised conscientiousness | not reported | .25 | .25 |
| Interests | .10 | .24 | −.08 |
| Overall conscientiousness | .31 | .21 | .02 |
| Unstructured interviews | .38 | .19 | −.01 |
| Years of job experience | .18 | .07 | −.07 |
The third column is the one most often ignored and the one Sackett and colleagues most want practitioners to use. It gives the tenth percentile of the distribution of validity values: roughly the worst case rather than the average. Read that way the ranking shifts again. Structured interviews have the higher mean, .42 against empirically keyed biodata's .38, but biodata has the better floor, .26 against .18. A risk-averse employer might reasonably prefer the second.
Two later revisions belong with the table. Sackett, Zhang, Berry and Lievens (2023) move assessment centres from .29 to .33 on a newer estimate of criterion reliability for managerial jobs, lifting them into a tie with work samples. In the other direction, Sackett, Demeke, Bazian, Griebie, Priest and Kuncel (2024) meta-analysed the relationship between cognitive ability and job performance across 153 twenty-first-century samples totalling 40,740 people, and found a mean observed validity of .16 and a corrected value of .22, which takes cognitive ability from fifth in the ranking to outside the top ten. They hypothesise that the decline reflects the shrinking share of manufacturing work and the growing role of team structures, which broaden what counts as performance beyond the task measures the older data relied on.
The rank order changed more than the individual figures did. Job-specific, sample-type methods rose to the top: structured interviews, job-knowledge tests, work samples, empirically keyed biodata and now assessment centres. General sign-type constructs, cognitive ability and broad personality among them, fell. One large drop has a separate explanation worth knowing. Schmidt and Hunter's .54 for work samples rested on seven primary studies from a 1984 meta-analysis; Roth, Bobko and McFarland (2005) added the studies conducted since and produced the .33 that Sackett et al. carry forward. That change has nothing to do with range restriction.
Two things that matter more than the coefficients
The dispute is live. Oh, Le and Roth (2023) argue that Sackett and colleagues overreach in recommending against range-restriction correction in concurrent designs, and identify conditions under which restriction is meaningful even there. Bobko, Roth, Le, Oh and Salgado (2025) go further, arguing that the ranking itself is compromised: where some predictors in the table are corrected for range restriction and others are not, a difference between two of them may reflect a true difference, differing degrees of restriction, differing correction, or some combination, and there is no way to tell which. They also object that the list compares predictor methods and predictor constructs side by side as though they were the same kind of thing. Their alternative is what they call considered estimation. Sackett, Berry, Lievens and Zhang (2025) have since replied, narrowing what they mean by conservative estimation: they endorse correction wherever the applicant-pool information needed to do it credibly exists, object only to defaulting to large corrections when it does not, and argue that most of Bobko and colleagues' objections are separate from that principle. Anyone presenting either set of figures as settled is misrepresenting the state of the field, and that includes anyone selling on the back of them.
The mean is not the number you will get. Sackett and colleagues introduced something the older syntheses omitted: standard deviations indicating between-study variability around each estimate. Those deviations differ sharply. Contextualised personality measures, meaning measures framed around behaviour at work rather than in general, show deviations near zero. Others are substantial: approximately .25 for interests, .20 for integrity tests, .19 for structured interviews, .17 for personality-based emotional intelligence, .16 for unstructured interviews and .15 for decontextualised conscientiousness. Structured interviews carry an eighty per cent credibility interval running from .18 to .66. A method with a good average can still perform poorly in a particular setting, and the averages alone will not tell you which setting you are in.
What the revised figures do not show
They do not show that validated instruments are worthless, and the authors are explicit that their work should not be read that way. Their own composite analysis makes the point. Across every possible combination of one to five predictors drawn from cognitive ability, conscientiousness, biodata, structured interviews and integrity tests, mean validity is .47 under the revised values against .51 under the prior ones. A six-predictor composite reaches .61.
That gap is the practical finding. Individual methods lost between .10 and .20 of their apparent validity. Batteries lost .04. The honest summary is that the effects are real and replicable, that they are considerably more modest per instrument than the field claimed for two decades, and that combining methods recovers most of the difference.
Part two: what the alternatives are worth
The counterfactual is not "nothing"
Organisations that do not use validated assessment are not making decisions in a vacuum. They are using something else, and the something else has been measured.
Unstructured interviews sit at .19 under the revised estimates, with a lower credibility value of .01. Years of job experience, the signal a CV carries most clearly, sits at .07, with a lower credibility value of .07. Years of education was not revisited by Sackett and colleagues; the Schmidt and Hunter figure of .10 remains the standing estimate. These are the three inputs on which a great many appointments are actually made.
Structure is the cheapest available improvement
The size of that gap is visible in the revised table itself: .42 against .19, the largest spread between two variants of a single method anywhere in it. Huffcutt and Arthur (1994) give the anatomy. Across 114 entry-level interview validity coefficients, classified on a multidimensional framework for level of structure, they found structure to be a major moderator of interview validity, and found that interviews, particularly structured ones, can reach validity comparable to that of mental ability tests.
Their third finding is the practically useful one and it is rarely quoted. Validity rises through much of the range of structure and then stops. Beyond a certain point, additional structure yields essentially no incremental validity. There is a ceiling. That matters for anyone designing a process, because the target is not maximum rigidity but enough structure to reach the plateau. Their analysis concerns entry-level jobs, which limits how far it generalises. McDaniel, Whetzel, Schmidt and Maurer (1994), across 245 coefficients and approximately 86,000 participants, found structured interviews clearly outperforming unstructured ones, with situational interviews higher still.
Structure means specific things: questions derived from a job analysis rather than invented in the room, the same questions in the same order for every candidate, anchored rating scales, independent scoring before discussion. It costs preparation time and nothing else. On the evidence it is the single highest-leverage change available to almost any employer, and it requires buying nothing.
Why experienced people resist it
Meehl's clinical-versus-statistical prediction tradition supplies part of the answer, and Grove and colleagues (2000) quantified it across 136 studies. Mechanical combination was about ten per cent more accurate than clinical judgement on average. It substantially outperformed clinical prediction in a third to a half of studies depending on the analysis, the two performed comparably in roughly half, and clinical judgement was substantially better in six to sixteen per cent. The advantage held regardless of the judgement task, the type of judge, or how experienced the judges were.
Two things about that deserve stating precisely, because it is routinely quoted more strongly than it warrants. Mechanical combination does not always win. And the meta-analysis covers human health and behaviour broadly rather than personnel selection specifically. What it establishes is a reliable modest advantage that survives expertise, which is a weaker claim than "the formula beats the person" and a more useful one.
The conclusion I draw is not that expertise is useless. It is that expertise needs feeding. A judgement built from good data points is a different thing from a judgement built from an impression, and neither is the same as a cut-off applied mechanically without anyone thinking about it. That is the argument for interpretation rather than automation, and it is why we put the weight we do on training people to read what an instrument actually says rather than to act on the number in front of them.
Highhouse (2008) offers an explanation, and identifies two implicit and rarely examined beliefs that inhibit the adoption of selection decision aids. They work in different ways.
The first is that near-perfect precision in predicting job performance is theoretically attainable. People resist analytical approaches, on this account, because they do not view selection as probabilistic and subject to error in the first place. A process designed as though prediction could be exact will feel more confident and perform no better.
The second he calls the myth of expertise: the belief that predicting human behaviour improves with experience. This produces an overreliance on intuition and, more awkwardly, a reluctance to use a decision aid at all, because reaching for one feels like conceding that your own judgement was not sufficient. That second mechanism is the difficult one, because it means the people most likely to resist structure are often the ones whose seniority makes their resistance decisive.
Kausel, Culbertson and Madrid (2016) demonstrated the mechanism experimentally. Giving decision makers unstructured interview information raised their confidence in their own judgements without improving accuracy, and under some conditions reduced it. A third study, a betting competition, found greater overconfidence associated with fewer payoffs, though that one used undergraduates rather than people who hire and is the weaker of the three legs. The first two carry the claim. This is the specific route by which adding unstructured interview data can degrade an otherwise reasonable decision. It is not that the extra information is neutral. It is that it makes people surer while making them no better.
Part three: where personality and trait EI actually sit
Conscientiousness and the contextualisation effect
Barrick and Mount (1991), across 117 studies, found conscientiousness the only Big Five trait showing consistent validity across all occupational groups and criteria, with extraversion showing validity in sales and management roles. Hurtz and Donovan (2000) refined the conscientiousness estimate to .20, the highest of the five. Sackett's revised figures put overall conscientiousness at .21 and overall emotional stability at .09.
The more useful recent finding concerns framing rather than trait. Sackett and colleagues show that contextualised personality measures, those asking about behaviour at work rather than in general, achieve higher validity and much lower between-study variance than decontextualised ones. Contextualised conscientiousness sits at .25 with a standard deviation of .00 and a lower credibility value of .25, meaning essentially every study in the distribution found the same thing. The general version sits at .21 with a standard deviation of .15 and a floor of .02. Contextualised emotional stability shows the same pattern, .23 against .09 for the general form. How you ask matters as much as what you ask about, and it matters most for consistency rather than for the average.
Emotional intelligence, and the distinction that does the work
The literature separates three streams, following Ashkanasy and Daus (2005). Stream one is ability EI, measured by maximum-performance tests such as the MSCEIT, where responses are scored right or wrong. Stream two is self-report or peer-report measurement of that same four-branch ability model. Stream three is everything else, usually described as mixed models of emotional competencies.
Where the TEIQue sits in that scheme is a real problem rather than a technicality. By default it falls into stream three, because it is not built on the four-branch ability model. But stream three is a residual category, and it groups the TEIQue with competency inventories built on entirely different foundations. Trait EI as Petrides defines it is neither an ability measure nor a competency inventory. It is a constellation of emotional self-perceptions, deliberately located within personality space rather than alongside cognitive ability.
Ability EI and trait EI are distinct constructs rather than two ways of measuring one thing, and the evidence bears that out. Their correlations with the general factor of personality are .28 and .88 respectively, and under the revised Sackett estimates ability-based EI shows validity for job performance of .22 against .30 for personality-based EI. Anyone treating figures reported for one as applying to the other is making a mistake the source data does not support.
O'Boyle, Humphrey, Pollack, Hawver and Story (2011) meta-analysed emotional intelligence and job performance across three streams: ability-based models using objective test items, self-report or peer-report measures based on the four-branch model, and mixed models of emotional competencies. All three showed corrected correlations with job performance in the range .24 to .30. Streams two and three added incremental validity beyond cognitive ability and the five-factor model, and dominance analysis found all three streams retained substantial relative importance alongside personality and intelligence.
One caveat travels with those figures and is rarely stated. That taxonomy defines its second stream by reference to the four-branch ability model, which is not the basis on which the TEIQue was constructed, so where exactly the TEIQue sits within it is a matter of genuine disagreement rather than settled classification. Figures reported for a stream do not transfer cleanly to any single instrument within or adjacent to it. Miao, Humphrey and Qian (2017) reported self-report EI correlating at .32 with job satisfaction, .43 with organisational commitment and negative .33 with turnover intention, and found it retained modest but statistically significant incremental validity in the presence of both cognitive ability and personality. Doğru (2022) reported, for self-report EI specifically, .33 with job performance, .37 with organisational citizenship behaviour, .31 with job satisfaction, .28 with organisational commitment and negative .45 with job stress.
On mechanism rather than magnitude, Miao and colleagues found the EI to job-satisfaction relationship mediated by state affect and by job performance itself, and Szczygiel and Mikolajczak (2018), in a study of nurses, found trait EI moderating the effect of anger and sadness at work on burnout: negative emotions were associated with greater burnout among nurses low in trait EI, but not among those high in it.
Two observations from the Sackett revision are worth stating together, because they cut in opposite directions. Personality-based EI at approximately .30 sits above both contextualised conscientiousness at .25 and overall conscientiousness at .21, which is notable given that conscientiousness has been treated as the standout personality predictor since Barrick and Mount. Against that, its between-study standard deviation is .17, among the larger ones in the table. The discipline applied to structured interviews above applies here without exception.
The redundancy question, and what actually answers it
Joseph, Jin, Newman and O'Boyle (2015) found that the content of mixed EI measures overlaps strongly with well-known constructs, including Big Five traits, cognitive ability and self-efficacy, and concluded that their predictive power is therefore partly borrowed. Because the TEIQue falls into that residual third stream by default, the critique is routinely read as applying to it.
It does not transfer cleanly, and the reason is the strongest thing in this section.
The critique bites against instruments that present themselves as measuring an intelligence while in fact measuring personality. Trait EI makes no such claim. It is defined as a personality construct from the outset. Overlap with the Big Five is therefore not an awkward discovery about it. It is what the theory predicts.
The scale of that overlap is now well documented and larger than most people expect. Van der Linden and colleagues (2017) meta-analysed the general factor of personality, the higher-order factor sitting above the Big Five, and found a true-score correlation with trait EI of .88, concluding the two may be virtually identical constructs. A subsequent twin study by the same group, with Petrides among the authors, found heritabilities of 53 per cent for the general factor and 45 per cent for trait EI, and a genetic correlation between them of .90. Ability EI, by contrast, correlated with the general factor at .28.
Two readings of that are available and the difference matters. One is that trait EI is redundant, measuring what personality inventories already capture. The other is that trait EI provides a direct measure of the apex of the personality hierarchy, which the Big Five capture only indirectly and which carries a substantive interpretation of its own: general social effectiveness, knowing what a situation calls for and being able to produce it.
Andrei and colleagues' incremental validity result, below, is what distinguishes between the two readings. If the overlap were complete there would be nothing left to add.
Two caveats belong here and I will not bury them. The existence of the general factor of personality is itself contested. And a construct correlating at .88 with something else is, on any reading, mostly that thing.
More broadly, Robinson and Zell (2026) synthesised 62 meta-analyses covering over 3,000 studies and around a million participants, reporting a robust association between an overall EI index and human flourishing at r = .28, with separate indices holding for ability, trait and mixed EI. That is a wide criterion rather than a work criterion. It does not describe a construct that dissolves under scrutiny.
Faking is the other standing objection, and the literature divides into two camps that have not reconciled.
The first holds that the effect on selection outcomes is minimal. Ones, Viswesvaran and Reiss (1996) meta-analysed the social-desirability literature and found that social-desirability scales did not predict school success, task performance, counterproductive behaviour or job performance, and that social desirability functioned neither as a predictor, nor as a practically useful suppressor, nor as a mediator for job performance. They also found it reflects real individual differences in emotional stability and conscientiousness rather than pure distortion, which is why they called it a red herring.
The second camp holds that faking degrades the properties that matter: raising means, compressing standard deviations, reducing reliability, and reducing criterion-related validity. Morgeson and colleagues (2007) sit outside both, and their position is worth reporting precisely because it is routinely mangled. Six former editors of Personnel Psychology and the Journal of Applied Psychology, with no commercial stake in testing, concluded that faking on self-report personality tests cannot be avoided and is perhaps not the issue at all; that the issue is instead the very low validity of such tests for predicting job performance; that using published self-report personality tests in selection should therefore be reconsidered; and that personality constructs may still have value, but that research should look for alternatives to self-report measurement.
That is a considerably stronger claim than the faking literature itself makes, and it drew two substantial rebuttals in the same journal, from Tett and Christiansen and from Ones and colleagues. It is a serious minority position rather than the settled view, and anyone citing it as though it closed the question is misusing it. It is also a challenge our own instrument family has to answer rather than wave away, the TEIQue being a self-report measure.
What the evidence does not support is treating the question as closed in either direction.
What can be said about the TEIQue specifically
Andrei, Siegling, Aloe, Baldaro and Petrides (2016) conducted the TEIQue-specific systematic review and meta-analysis. Reviewing 24 articles reporting 114 incremental-validity analyses, and pooling 18 studies providing 105 effect sizes, they found the TEIQue explained incremental variance beyond higher-order personality in 84.2% of analyses, with a pooled effect size of ΔR² = .06 (95% CI .03 to .08), driven mainly by the well-being and self-control factors. Their own bias checks add a caveat that belongs with the figure: Egger's test and the funnel plot showed statistically significant asymmetry, so the pooled effect may overestimate the underlying one.
Three things belong alongside that. The criteria in that review span affect, behaviour, cognition, desire and somatic health, which is to say they are largely not work criteria. The meta-analytic work-outcome evidence covers self-report EI as a family rather than the TEIQue in isolation. And the workplace evidence that does exist is overwhelmingly cross-sectional: Li, Pérez-Díaz, Mao and Petrides (2018), a multilevel study of 881 Chinese primary-school teachers in which job satisfaction partially mediated a positive association between trait EI and job performance, is among the strongest of it, and it is a snapshot rather than a prediction across time.
The defensible claim is incremental validity, not dominance. These are associations, not demonstrated causal effects, and the word "predicts" should be reserved for designs that actually predict. Longitudinal and experimental selection studies remain scarce, and that is a real evidence gap rather than a presentational inconvenience.
A limit follows directly. No instrument of ours, and none of anyone else's, should serve as the sole basis for an employment decision. That is not a compliance disclaimer appended at the end. It is a statement about what a correlation of around .30 can and cannot be asked to carry.
The same standard disposes of the tools most commonly reached for instead, and it is the publishers themselves who say so most plainly. The Myers and Briggs Foundation's ethical guidelines state that it is not ethical to use the MBTI for hiring or for deciding job assignments. A leading DISC publisher states that DISC is not recommended for pre-employment screening because it does not measure any skill, aptitude or factor specific to a position, and that it is not a predictive assessment. Neither instrument appears anywhere in the Schmidt and Hunter hierarchy or in the Sackett revision, and the publishers' own statements above are the plainest available explanation of why. Pittenger's (2005) review reached the same conclusion, and reported that around half of respondents receive a different four-letter type on retest.
Represented fairly, most serious critiques of personality testing in hiring turn out to be critiques of misused typological tools and of overclaiming, rather than of validated dimensional measurement.
Part four: the cost of getting it wrong
Why I will not open with a dramatic figure
Our category has a habit of leading with a headline cost for a bad hire. Having looked at where those figures come from, I will not be doing it.
The best-sourced general benchmark is Boushey and Glynn (2012), who reviewed thirty case studies drawn from eleven research papers published between 1992 and 2007. The circulating alternatives are worth setting beside it, because they differ enormously in what stands behind them.
| Figure | Source | What it rests on |
|---|---|---|
| ≈ 20% of annual salary | Boushey & Glynn (2012) | 22 case studies covering jobs paying under $50,000, roughly three quarters of US workers |
| 21% median | Boushey & Glynn (2012) | 27 case studies, excluding executives and physicians |
| Up to 213% | Boushey & Glynn (2012) | Very high-skill roles only, where replacement is genuinely hard |
| 50% to 200% | Professional-body benchmarking | Widely circulated, methodology rarely transparent |
| ≥ 30% of first-year salary | Attributed to US Dept of Labor | Cited almost everywhere; I could not trace it to a primary publication with a stated method |
The pattern is worth noticing. The best-documented figure is the least dramatic, and the dramatic ones are the least documented.
On downside risk specifically, Housman and Minor (2015) analysed 58,542 workers across eleven firms and estimated that avoiding a toxic worker was associated with roughly $12,489 in avoided induced turnover costs, against approximately $5,303 for the gain from a top one per cent performer, falling to roughly $1,951 for a top quartile performer. The authors state that the $12,489 figure excludes litigation, regulatory penalty and reduced morale, so it is a floor rather than a total.
The scope matters and is easy to overlook. It is a working paper on observational data, the relationship is correlational, all the employees studied were in hourly front-line service roles, and toxicity was defined narrowly as termination for serious misconduct, covering around five per cent of workers. Read within those limits it still points somewhere counter-intuitive: avoiding the worst hire may matter more than securing the best one.
The finding our industry should sit with
Latham and Whyte (1994), in a paper titled "The futility of utility analysis", studied 143 experienced managers and found that utility analysis influenced their decisions, but not in the direction its advocates intended. Presenting the figures reduced managerial support for implementing a valid selection procedure, even though the analysis showed the net benefits were substantial.
The finding did not go unchallenged. Carson, Becker and Henderson (1998) failed to replicate it across two studies of 145 and 186 participants, and found that when utility information was presented in a revised, clearer format, it had a low-to-moderate positive effect on managers' acceptance of the same kind of proposal. Read together, the studies support a narrower claim than either does alone: financial evidence about selection does not speak for itself, and how it is presented, and by whom, shapes whether it persuades.
The boundary condition makes this worse for our industry rather than better. Cronshaw (1997) argued the effect holds specifically where the presenter is perceived as selling an intervention, rather than advising at arm's length. Which is to say it applies most precisely to the situation where a psychometrics company puts a cost-of-bad-hire figure in front of a prospective client.
Replicated or not in any given room, I find that possibility harder to dismiss than any of the validity coefficients above.
Part five: what the reading changed
Working through this research surfaced something more specific than I had expected, and it turned out not to be about instruments alone.
The largest single gap in the revised table sits between two variants of one method: the structured interview at .42 and the unstructured conversation at .19. Most real interviews sit somewhere between those poles rather than at either end, and what separates them is not talent or seniority but preparation: questions derived from the role, the same questions in the same order for every candidate, anchored rating scales, independent scoring before discussion.
Whatever an interview does not settle does not disappear. It moves. It is paid for later, in training, debriefing and review, and spotting a mismatch early is worth considerably more than spotting it late. The interview is where early happens.
So the conclusion is not simply that an instrument should move earlier in the sequence, though it should. It is that selection should be designed as a system from the outset: a structured interview, with a validated assessment used as one of the selection criteria rather than as an induction exercise. That combination does two things worth having. It gives the interview something firmer to work from. And it can produce a well-supported rejection, which is a far better use of resources than months spent training someone who was never going to fit the role.
There is a further reason to build both rather than either, and it is the composite finding above applied to a single role. Methods combine better than they perform alone, because each carries information the others do not. A structured interview and a validated assessment are close to the ideal version of that pairing. One samples reasoning and behaviour against criteria defined in advance. The other measures dispositions an interview cannot observe directly. Used together they should be more accurate than either on its own, and considerably more accurate than an unstructured conversation followed by hoping for the best.
That is the central thing I take from this research: the weight of evidence behind structured interviewing, and how cheap it is relative to what it saves. It costs preparation rather than money, and it is far easier to design in from the beginning than to retrofit later. How to build that structure well, and how a validated assessment sits alongside a structured interview without either being asked to do the other's job, is a subject in its own right.
What I take from it
One qualification before the general point. Everything above is framed as systems design, and hiring only partly is. Criteria, scoring and records are systems. The people on either side of the table are not. Structure is there to make judgement about people more consistent and more defensible, not to replace it, and any reading of this evidence that ends in a formula deciding who gets hired has gone wrong somewhere. The instruments inform the decision. They do not make it.
The generalisable part is not about psychometrics at all.
Leaving your own comfort zone and taking on responsibilities outside it brings a benefit that is easy to underrate. It gives you a second angle on your own work. Sitting with a body of evidence for one purpose and then turning to look at what you have built with that evidence still in view is a genuinely different vantage point, and it shows you things that reviewing your own work on its own terms never will. I would not have caught this from inside the operational role alone. I caught it because I was also the person doing the reading.
That is worth something to anyone, in any field. The useful question is not whether you believe the evidence on structured selection. It is whether you have looked at your own work recently from somewhere other than where you usually stand.
A note on sources
Figures here are drawn from a research review conducted for the London Psychometric Laboratory and are labelled by source type where that matters. Meta-analytic validity coefficients come from peer-reviewed journals. Cost-of-turnover figures come from a mixture of independent research synthesis, professional-body benchmarking and one large observational working paper, which differ substantially in evidential weight. All trait EI findings cited here are correlational and predominantly cross-sectional. Where a widely circulated figure could not be traced to a primary source with a stated methodology, that is stated rather than omitted.
References
Andrei, F., Siegling, A. B., Aloe, A. M., Baldaro, B., & Petrides, K. V. (2016). The incremental validity of the Trait Emotional Intelligence Questionnaire (TEIQue): A systematic review and meta-analysis. Journal of Personality Assessment, 98(3), 261–276. https://doi.org/10.1080/00223891.2015.1084630
Ashkanasy, N. M., & Daus, C. S. (2005). Rumors of the death of emotional intelligence in organizational behavior are vastly exaggerated. Journal of Organizational Behavior, 26(4), 441–452.
Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26. https://doi.org/10.1111/j.1744-6570.1991.tb00688.x
Bobko, P., Roth, P. L., Le, H., Oh, I.-S., & Salgado, J. F. (2025). The need for "considered estimation" versus "conservative estimation" when ranking or comparing predictors of job performance. International Journal of Selection and Assessment, 33(1).
Boushey, H., & Glynn, S. J. (2012). There are significant business costs to replacing employees. Center for American Progress.
Carson, K. P., Becker, J. S., & Henderson, J. A. (1998). Is utility really futile? A failure to replicate and an extension. Journal of Applied Psychology, 83(1), 84–96. https://doi.org/10.1037/0021-9010.83.1.84
Cronshaw, S. F. (1997). Lo! The stimulus speaks: The insider's view on Whyte and Latham's "The futility of utility analysis". Personnel Psychology, 50, 611–616.
Doğru, Ç. (2022). A meta-analysis of the relationships between emotional intelligence and employee outcomes. Frontiers in Psychology, 13, Article 611348. https://doi.org/10.3389/fpsyg.2022.611348
Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., & Nelson, C. (2000). Clinical versus mechanical prediction: A meta-analysis. Psychological Assessment, 12(1), 19–30.
Highhouse, S. (2008). Stubborn reliance on intuition and subjectivity in employee selection. Industrial and Organizational Psychology, 1(3), 333–342. https://doi.org/10.1111/j.1754-9434.2008.00058.x
Housman, M., & Minor, D. (2015). Toxic workers. Harvard Business School Working Paper 16-057.
Huffcutt, A. I., & Arthur, W. (1994). Hunter and Hunter (1984) revisited: Interview validity for entry-level jobs. Journal of Applied Psychology, 79(2), 184–190.
Hurtz, G. M., & Donovan, J. J. (2000). Personality and job performance: The Big Five revisited. Journal of Applied Psychology, 85(6), 869–879.
Joseph, D. L., Jin, J., Newman, D. A., & O'Boyle, E. H. (2015). Why does self-reported emotional intelligence predict job performance? A meta-analytic investigation of mixed EI. Journal of Applied Psychology, 100(2), 298–342.
Kausel, E. E., Culbertson, S. S., & Madrid, H. P. (2016). Overconfidence in personnel selection: When and why unstructured interview information can hurt hiring decisions. Organizational Behavior and Human Decision Processes, 137, 27–44.
Latham, G. P., & Whyte, G. (1994). The futility of utility analysis. Personnel Psychology, 47(1), 31–46. https://doi.org/10.1111/j.1744-6570.1994.tb02408.x
Li, M., Pérez-Díaz, P. A., Mao, Y., & Petrides, K. V. (2018). A multilevel model of teachers' job performance: Understanding the effects of trait emotional intelligence, job satisfaction, and organizational trust. Frontiers in Psychology, 9, 2420. https://doi.org/10.3389/fpsyg.2018.02420
McDaniel, M. A., Whetzel, D. L., Schmidt, F. L., & Maurer, S. D. (1994). The validity of employment interviews: A comprehensive review and meta-analysis. Journal of Applied Psychology, 79(4), 599–616.
Miao, C., Humphrey, R. H., & Qian, S. (2017). A meta-analysis of emotional intelligence and work attitudes. Journal of Occupational and Organizational Psychology, 90(2), 177–202. https://doi.org/10.1111/joop.12167
Morgeson, F. P., Campion, M. A., Dipboye, R. L., Hollenbeck, J. R., Murphy, K., & Schmitt, N. (2007). Reconsidering the use of personality tests in personnel selection contexts. Personnel Psychology, 60(3), 683–729.
O'Boyle, E. H., Humphrey, R. H., Pollack, J. M., Hawver, T. H., & Story, P. A. (2011). The relation between emotional intelligence and job performance: A meta-analysis. Journal of Organizational Behavior, 32(5), 788–818. https://doi.org/10.1002/job.714
Oh, I.-S., Le, H., & Roth, P. L. (2023). Revisiting Sackett et al.'s (2022) rationale behind their recommendation against correcting for range restriction in concurrent validation studies. Journal of Applied Psychology, 108(8), 1300–1310. https://doi.org/10.1037/apl0001078
Ones, D. S., Dilchert, S., Viswesvaran, C., & Judge, T. A. (2007). In support of personality assessment in organizational settings. Personnel Psychology, 60(4), 995–1027.
Ones, D. S., Viswesvaran, C., & Reiss, A. D. (1996). Role of social desirability in personality testing for personnel selection: The red herring. Journal of Applied Psychology, 81(6), 660–679.
Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221. https://doi.org/10.1037/1065-9293.57.3.210
Robinson, T. J., & Zell, E. (2026). Robust associations of emotional intelligence with human flourishing: A second-order meta-analysis. Proceedings of the National Academy of Sciences, 123, e2532963123. https://doi.org/10.1073/pnas.2532963123
Roth, P. L., Bobko, P., & McFarland, L. A. (2005). A meta-analysis of work sample test validity: Updating and integrating some classic literature. Personnel Psychology, 58(4), 1009–1037.
Sackett, P. R., Berry, C. M., Lievens, F., & Zhang, C. (2025). Issues in contrasting conservative and considered estimation: A reply to Bobko et al. (2024). International Journal of Selection and Assessment, 33(3), e70016. https://doi.org/10.1111/ijsa.70016
Sackett, P. R., Demeke, S., Bazian, I. M., Griebie, A. M., Priest, R., & Kuncel, N. R. (2024). A contemporary look at the relationship between general cognitive ability and job performance. Journal of Applied Psychology, 109(5), 687–713. https://doi.org/10.1037/apl0001159
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000994
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, 16(3), 283–300. https://doi.org/10.1017/iop.2023.24
Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.
Szczygiel, D. D., & Mikolajczak, M. (2018). Emotional intelligence buffers the effects of negative emotions on job burnout in nursing. Frontiers in Psychology, 9, Article 2649. https://doi.org/10.3389/fpsyg.2018.02649
Tett, R. P., & Christiansen, N. D. (2007). Personality tests at the crossroads: A response to Morgeson, Campion, Dipboye, Hollenbeck, Murphy, and Schmitt (2007). Personnel Psychology, 60(4), 967–993.
Van der Linden, D., Pekaar, K. A., Bakker, A. B., Schermer, J. A., Vernon, P. A., Dunkel, C. S., & Petrides, K. V. (2017). Overlap between the general factor of personality and emotional intelligence: A meta-analysis. Psychological Bulletin, 143(1), 36–52.
Van der Linden, D., Schermer, J. A., de Zeeuw, E., Dunkel, C. S., Pekaar, K. A., Bakker, A. B., Vernon, P. A., & Petrides, K. V. (2018). Overlap between the general factor of personality and trait emotional intelligence: A genetic correlation study. Behavior Genetics, 48(2), 147–154.
