Job descriptions keep growing because companies buy people when what they have is a gap, and a role is the only unit most organisations know how to purchase. The bundle costs candidates in self-disqualification and stress, and costs companies in mediocre coverage across two thirds of a role they paid full price for. Meanwhile the interview format used to select for that bundle has been measured for decades, and the structured version outperforms it in every major analysis. AI has changed the volume on both sides of this process and has not touched its accuracy. It also, in the same period, made the high-validity methods affordable for the first time, and almost nobody has spent the windfall on those.
The complaint, stated precisely
Anyone who has applied for a senior role recently has noticed the same three things. The job description lists more requirements than any one person plausibly holds. The interview tests whether you can recall a framework and repeat it under time pressure. And somewhere in the middle a hypothetical problem arrives, built on data you cannot see, to be solved in four minutes without the thinking time that any real version of that problem would obviously get.
The usual reading of this is that the candidate is not a fit, or is bad at interviews. There is a more useful reading available, which is that the unstructured interview predicts performance worse than a structured one, which has been published since the 1998 meta-analysis and holds in its 2022 revision.
What the selection research actually found
Personnel selection is one of the most heavily measured areas in applied psychology, because the outcome is observable and the sample sizes are enormous. The foundational work is Schmidt and Hunter's 1998 meta-analysis of 85 years of research across nineteen selection methods. In 2022, Sackett, Zhang, Berry and Lievens showed that the 1998 corrections for range restriction were too aggressive and re-estimated the validities, and their figures are the current best estimate. Each method carries an operational validity, a correlation with subsequent job performance where 0 is a coin flip and 1 is perfect prediction.
Operational validity: 2022 revision and 1998 estimate
Sackett, Zhang, Berry and Lievens (2022), Journal of Applied Psychology, as summarised by SIOP; Schmidt and Hunter (1998), Psychological Bulletin. Operational validity. The 2022 paper also lowers the unstructured interview estimate, so the 1998 figure of .38 is an upper bound.
Two things fall out of that table. Every estimate fell in 2022, and the structured interview came out as the strongest single predictor, ahead of general mental ability and work samples. And the unstructured interview, which was already the weakest of the interview formats in 1998, sits below the structured one in both analyses, so the practical advice has not changed even though the numbers have. Work samples, an actual piece of the actual job done under conditions resembling the actual job, remain a solid predictor at .33 on the 2022 figures.
Now compare that to what the process usually looks like. The ambiguous question where the candidate is guessing at an unstated preference is, definitionally, an unstructured interview, the format that ranks below the structured one in both analyses. The hypothetical problem on invisible data is a work sample with the work removed, which retains the stress of the format and discards the thing that made it predictive. And the home assignment, the group exercise, and the demonstrated ability to learn something unfamiliar, which sit closest to genuine work samples, are the parts most often treated as optional extras or skipped for speed.
Why the description kept growing
The requirements inflation is a separate failure with a shared root, and it is the one I find more interesting because it is entirely self-inflicted.
A company rarely has a person-shaped hole. It has a set of things that are not getting done: a migration nobody owns, a report that takes three days a month, a product area without a decision-maker, an onboarding process that leaks. Those are gaps, and they arrive at different times, in different sizes, and with different half-lives.
But a company cannot purchase a gap. It can purchase a role. So the gaps get bundled into a single document, and because nobody wants to reopen the requisition in six months, everything anyone might conceivably need gets added while the document is open. The result is a description that is not a description of a job. It is a wish list assembled by committee, standing in for a specification nobody wrote.
This is the same failure the GEAR model addresses from the organisational side: the org chart optimises for what things are called rather than what the organisation actually needs. Hiring is where that abstraction gets expensive, because the bundle is now attached to a salary.
What the bundle costs, on both sides
| What happens | What it costs | |
|---|---|---|
| The candidate | Reads twelve requirements, holds eight of them well, self-disqualifies or applies while braced for rejection | Strong people filter themselves out before anyone assesses them, and the ones who apply arrive carrying the belief that they are already short |
| The hire | Lands in a role built from five jobs, is genuinely strong at two of them | Excellent output in a third of the role, adequate in a third, and quiet failure in the rest, which nobody planned for and nobody is measuring |
| The company | Concludes the person is underperforming rather than that the role was unbuildable | A second hire to cover the weak third. The bundle has now produced two salaries where the actual demand was one and a half |
| The organisation | Repeats this per department, per year | Permanent headcount accumulated from temporary gaps, with the loaded cost of every seat carried indefinitely |
That third row is the expensive one and it is almost never diagnosed correctly. When someone is strong in part of a bundled role and weak elsewhere, the organisation experiences it as a performance problem, because performance problems have a familiar shape and role-design problems do not. The remedy applied is another person, which means the original bundling error has now been converted into permanent cost.
There is also a quieter effect worth naming. A person who spends two years being adequate at a third of their job, in an environment that treats that as a shortfall rather than as arithmetic, does not conclude that the role was badly specified. They conclude something about themselves. That belief follows them into the next interview, where it reads as diminished confidence, which the unstructured interview is exquisitely sensitive to and terrible at interpreting.
Did AI change any of this
It changed the volume enormously and the accuracy not at all, which is the worst available combination.
On the candidate side, AI-assisted applications are now common: CNBC reported in February 2025 on a Career Group Companies survey finding that nearly two thirds of job candidates use AI somewhere in their applications. On the employer side, screening has moved further into automated parsing and semantic scoring, and in a Resume.org survey of nearly 1,400 US workers reported by HR Dive in August 2025, a third said AI would likely run their company's entire hiring process by 2026. Both figures come from commercial surveys rather than official statistics and should be read as directional.
The direction, though, is not in doubt. More applications are produced with less effort, so more filtering is applied earlier, so applications are optimised for the filter rather than for the reader, so the filter is tightened. Everyone is now moving faster through a process whose predictive core, the interview itself, is unchanged.
The part nobody is using
There is a second half to what AI did here, and it is the more interesting one, because it went almost entirely unclaimed.
The honest reason most companies never ran the high-validity methods was not ignorance. It was cost. Writing a structured interview kit means deciding what the seat is actually for, drafting a question set, building a scoring rubric, and training four interviewers to use it identically. Building a realistic work sample means constructing a task on data that resembles the real thing, with an answer key someone can mark against. Per role, that was, in my experience, weeks of specialist work, which is why it was reserved for volume hiring at large companies and skipped everywhere else. The evidence for structure was never the obstacle. The economics were.
That constraint is gone. A scorecard, a structured question set with a rubric, and a work sample built on synthetic-but-realistic data are now, in my experience, an afternoon of work for anyone who knows what the seat needs to produce. The methods with the best evidence behind them became affordable for a fifteen-person company in the same year that title-keyword screening became automatic.
So the choice available in 2026 is not between an expensive good process and a cheap bad one. Both are cheap now. Many companies spent the windfall on running the unstructured process faster, because that was the part already sitting inside the applicant tracking system, and the structured version still requires someone to sit down and decide what the role is for. The constraint stopped being cost and became attention, and the spending pattern has not caught up.
What to do instead
The interventions with the best evidence are unglamorous, and none of them require new technology.
Structure the interview. Same questions, same order, same scoring rubric, scored independently before discussion. On the 1998 estimates this moved the format from .38 to .51; the 2022 revision lowers both figures and still places the structured interview at .42, above the unstructured one. It is the cheapest improvement available in the process and takes an afternoon of preparation per role.
Use a real work sample. A genuine slice of the job, on real or realistic data, with the thinking time the job would actually allow. If the role involves building a report, have them build a small report. The hypothetical-under-pressure format is not a lightweight version of this; it measures something else.
Separate the must-haves from the wish list before the description is written. A useful discipline: name the three things this person will be measured on in year one. Anything not serving those three is a preference, and preferences belong at the bottom of the document clearly labelled as such.
Ask whether the gap needs a person at all. Some gaps are permanent and deserve a permanent seat. Many are a defined piece of work with an end date, and a role is a very expensive way to buy those, because the salary continues after the gap closes.
The connection to how I work
This is the reasoning behind how I sell. What a company usually needs is not a person; it is a migration finished, a report built, a role covered for six months, a system made to do what it was bought for. Those are gaps, and buying a permanent seat to close a temporary gap is how organisations end up carrying headcount that outlived its reason.
So the unit I sell is the gap rather than the person: a defined piece of work, paid for once. It is a smaller commercial unit than a hire, which is precisely the point. The jobs people actually call about are bundle components that were never worth a whole role.
Sources
- Schmidt, F. L. and Hunter, J. E., "The Validity and Utility of Selection Methods in Personnel Psychology: Practical and Theoretical Implications of 85 Years of Research Findings", Psychological Bulletin, 1998. The 1998 meta-analysis across 19 selection methods; source of the older estimates (structured interview .51, general mental ability .51, work sample .54, unstructured interview .38).
- Sackett, P. R., Zhang, C., Berry, C. M. and Lievens, F., "Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range", Journal of Applied Psychology, 107(11), 2040-2068, 2022. Source of the revised estimates.
- Society for Industrial and Organizational Psychology, "Is Cognitive Ability the Best Predictor of Job Performance? New Research Says It's Time to Think Again". Summary of the 2022 values quoted above (structured interview .42, general mental ability .31, work sample .33).
- CNBC, "Nearly two-thirds of job candidates are using AI in their applications, report says", 28 February 2025, on a Career Group Companies report. Industry survey, directional.
- HR Dive, "1 in 3 companies say AI will run their hiring process by 2026", August 2025, on a Resume.org survey of nearly 1,400 US workers. Industry survey, directional.
- Companion piece: The GEAR Model, on designing the organisation around what it needs rather than what roles are called.