Insights / Organizational design

The job description is a bundle

In short

Job descriptions keep growing because companies buy people when what they have is a gap, and a role is the only unit most organisations know how to purchase. The bundle costs candidates in self-disqualification and stress, and costs companies in mediocre coverage across two thirds of a role they paid full price for. Meanwhile the interview format used to select for that bundle has been measured for a century and performs near the bottom of every table. AI has changed the volume on both sides of this process and has not touched its accuracy. It also, in the same period, made the high-validity methods affordable for the first time, and almost nobody has spent the windfall on those.

The complaint, stated precisely

Anyone who has applied for a senior role recently has noticed the same three things. The job description lists more requirements than any one person plausibly holds. The interview tests whether you can recall a framework and repeat it under time pressure. And somewhere in the middle a hypothetical problem arrives, built on data you cannot see, to be solved in four minutes without the thinking time that any real version of that problem would obviously get.

The usual reading of this is that the candidate is not a fit, or is bad at interviews. There is a more useful reading available, which is that the process is measurably worse at predicting performance than the alternatives, and this has been known and published since 1998.

What a century of research actually found

Personnel selection is one of the most heavily measured areas in applied psychology, because the outcome is observable and the sample sizes are enormous. The foundational work is Schmidt and Hunter's 1998 meta-analysis of 85 years of research across nineteen selection methods, updated by Schmidt, Oh and Shaffer in 2016 to cover a century. Each method carries an operational validity, a correlation with subsequent job performance where 0 is a coin flip and 1 is perfect prediction.

Operational validity, selected methods

GMA + integrity test.65
GMA + work sample.63
GMA + structured interview.63
Structured interview alone.51
General mental ability alone.51
Unstructured interview.38

Schmidt and Hunter (1998), Psychological Bulletin; Schmidt, Oh and Shaffer (2016) update. Operational validity, corrected for range restriction.

Two things fall out of that table immediately. The structured interview is one of the two best single predictors available, and the unstructured interview is one of the weakest things a company can do while still believing it is running a selection process. And the highest-performing combinations all involve a work sample: an actual piece of the actual job, done under conditions resembling the actual job.

Now compare that to what the process usually looks like. The ambiguous question where the candidate is guessing at an unstated preference is, definitionally, an unstructured interview. It is the .38. The hypothetical problem on invisible data is a work sample with the work removed, which retains the stress of the format and discards the thing that made it predictive. And the home assignment, the group exercise, and the demonstrated ability to learn something unfamiliar, which sit closest to genuine work samples, are the parts most often treated as optional extras or skipped for speed.

The methods with the strongest evidence behind them are the ones companies use least, and the method with the weakest evidence is the one almost every process is built around.

Why the description kept growing

The requirements inflation is a separate failure with a shared root, and it is the one I find more interesting because it is entirely self-inflicted.

A company rarely has a person-shaped hole. It has a set of things that are not getting done: a migration nobody owns, a report that takes three days a month, a product area without a decision-maker, an onboarding process that leaks. Those are gaps, and they arrive at different times, in different sizes, and with different half-lives.

But a company cannot purchase a gap. It can purchase a role. So the gaps get bundled into a single document, and because nobody wants to reopen the requisition in six months, everything anyone might conceivably need gets added while the document is open. The result is a description that is not a description of a job. It is a wish list assembled by committee, standing in for a specification nobody wrote.

This is the same failure the GEAR model addresses from the organisational side: the org chart optimises for what things are called rather than what the organisation actually needs. Hiring is where that abstraction gets expensive, because the bundle is now attached to a salary.

What the bundle costs, on both sides

What happensWhat it costs
The candidate Reads twelve requirements, holds eight of them well, self-disqualifies or applies while braced for rejection Strong people filter themselves out before anyone assesses them, and the ones who apply arrive carrying the belief that they are already short
The hire Lands in a role built from five jobs, is genuinely strong at two of them Excellent output in a third of the role, adequate in a third, and quiet failure in the rest, which nobody planned for and nobody is measuring
The company Concludes the person is underperforming rather than that the role was unbuildable A second hire to cover the weak third. The bundle has now produced two salaries where the actual demand was one and a half
The organisation Repeats this per department, per year Permanent headcount accumulated from temporary gaps, with the loaded cost of every seat carried indefinitely

That third row is the expensive one and it is almost never diagnosed correctly. When someone is strong in part of a bundled role and weak elsewhere, the organisation experiences it as a performance problem, because performance problems have a familiar shape and role-design problems do not. The remedy applied is another person, which means the original bundling error has now been converted into permanent cost.

There is also a quieter effect worth naming. A person who spends two years being adequate at a third of their job, in an environment that treats that as a shortfall rather than as arithmetic, does not conclude that the role was badly specified. They conclude something about themselves. That belief follows them into the next interview, where it reads as diminished confidence, which the unstructured interview is exquisitely sensitive to and terrible at interpreting.

Did AI change any of this

It changed the volume enormously and the accuracy not at all, which is the worst available combination.

On the candidate side, AI-assisted applications have risen sharply, with commercial trackers putting usage somewhere near 30% of applicants during 2025, up from roughly half that a year earlier. On the employer side, screening has moved further into automated parsing and semantic scoring, and a substantial share of companies now expect AI to run most of the hiring pipeline. Both figures come from industry trackers rather than official statistics and should be read as directional.

The direction, though, is not in doubt. More applications are produced with less effort, so more filtering is applied earlier, so applications are optimised for the filter rather than for the reader, so the filter is tightened. Everyone is now moving faster through a process whose predictive core, the interview itself, is unchanged.

The point that matters for a CEO: AI has made a low-validity process cheaper to run at scale. Cheaper and faster are improvements when the underlying method works. When the method predicts performance at .38, volume is not the constraint that was worth relieving.

The part nobody is using

There is a second half to what AI did here, and it is the more interesting one, because it went almost entirely unclaimed.

The honest reason most companies never ran the high-validity methods was not ignorance. It was cost. Writing a structured interview kit means deciding what the seat is actually for, drafting a question set, building a scoring rubric, and training four interviewers to use it identically. Building a realistic work sample means constructing a task on data that resembles the real thing, with an answer key someone can mark against. Per role, that was weeks of specialist work, which is why it was reserved for volume hiring at large companies and skipped everywhere else. The evidence was never in dispute. The economics were.

That constraint is gone. A scorecard, a structured question set with a rubric, and a work sample built on synthetic-but-realistic data are now an afternoon of work for anyone who knows what the seat needs to produce. The methods that predict performance at .51 and .63 became affordable for a fifteen-person company in the same year that title-keyword screening became automatic.

So the choice available in 2026 is not between an expensive good process and a cheap bad one. Both are cheap now. Companies overwhelmingly spent the windfall on running the .38 faster, because that was the part already sitting inside the applicant tracking system, and the .51 still requires someone to sit down and decide what the role is for. The constraint stopped being cost and became attention, and the spending pattern has not caught up.

What to do instead

The interventions with the best evidence are unglamorous, and none of them require new technology.

Structure the interview. Same questions, same order, same scoring rubric, scored independently before discussion. This alone moves the format from .38 to .51, which is the largest single improvement available in the whole process, and it costs one afternoon of preparation per role.

Use a real work sample. A genuine slice of the job, on real or realistic data, with the thinking time the job would actually allow. If the role involves building a report, have them build a small report. The hypothetical-under-pressure format is not a lightweight version of this; it measures something else and predicts less.

Separate the must-haves from the wish list before the description is written. A useful discipline: name the three things this person will be measured on in year one. Anything not serving those three is a preference, and preferences belong at the bottom of the document clearly labelled as such.

Ask whether the gap needs a person at all. Some gaps are permanent and deserve a permanent seat. Many are a defined piece of work with an end date, and a role is a very expensive way to buy those, because the salary continues after the gap closes.

The connection to how I work

This is the reasoning behind how I sell. What a company usually needs is not a person; it is a migration finished, a report built, a role covered for six months, a system made to do what it was bought for. Those are gaps, and buying a permanent seat to close a temporary gap is how organisations end up carrying headcount that outlived its reason.

So the unit I sell is the gap rather than the person: a defined piece of work, paid for once, ending in a written handover so the same task does not come back to me. It is a smaller commercial unit than a hire, which is precisely the point. The nine jobs people actually call about are all bundle components that were never worth a whole role.

Sources

  • Schmidt, F. L. and Hunter, J. E., "The Validity and Utility of Selection Methods in Personnel Psychology: Practical and Theoretical Implications of 85 Years of Research Findings", Psychological Bulletin, 1998. The foundational meta-analysis across 19 selection methods.
  • Schmidt, F. L., Oh, I.-S. and Shaffer, J. A., the 2016 update covering 100 years of research findings. Source of the combined-validity figures for general mental ability paired with work samples, integrity tests, and structured interviews.
  • AI adoption figures on both sides of the hiring process come from commercial recruitment-technology trackers rather than official statistics, and are described above as directional rather than precise.
  • Companion piece: The GEAR Model, on designing the organisation around what it needs rather than what roles are called.
The Operating Baseline Which of your open roles are gaps, and which are genuinely seats
Book a 30-min call Tell me what is open right now. I answer 7 days a week.