KPMG and the McCombs School of Business at UT Austin gave 523 early-career professionals a set of realistic business tasks and a domain-specific AI agent, then had the AI complete the same tasks alone as a benchmark. Three groups emerged, which the researchers named Amplifiers, Delegators, and Apprentices. Amplifiers beat the AI. Delegators matched it. Apprentices came in below it, meaning their involvement made the work worse than leaving the model to it. The foundational capabilities most hiring and training programs target, critical thinking, domain knowledge, and AI literacy, did not predict which group anyone landed in. Apprentices scored as well on the fundamentals as Amplifiers. They questioned the model's output often, and their questions rarely improved it.
That is a problem for the way most companies are currently spending. A license for everyone and a prompt engineering course feel like progress because they are countable, and neither one touches what separated the groups, which was how people structured the problem, tested what came back, and folded it into their own judgment. BCG's survey of nearly 12,000 workers found the same split at the company level: clear direction on where AI is going and how to use it raised measurable business impact by 25 percentage points, while better tools raised it by five.
Delegators are the hardest group to see. Someone who accepts what the model produces meets deadlines and turns in acceptable work, so nothing in an adoption dashboard looks wrong. Seats, logins and prompt counts record whether people are using AI, not whether they added anything to it. The study could tell the difference because it compared each person's output against the same agent working alone, which is a comparison almost no company runs.
Five places to start
- Pick one workflow that already has a number on it, and redesign the whole loop. Choose a process you already report on, time to fill, comp cycle turnaround, onboarding completion, so the before and after is a figure finance already recognizes. Redesign the steps around AI rather than inserting AI into the existing steps. The pilots that produced measurable returns were narrow rather than broad.
- Find your Amplifiers and make their work the curriculum. Two or three people per function are already getting better results than the rest. Sit with them and record a real session: the opening prompt, the point where they pushed back, what they rewrote. That library becomes your training. A generic AI course teaches the tool, and tool knowledge is not what the study found separated people.
- Require one line of reasoning on AI-assisted work. When AI contributes to something that reaches a client, a candidate, or an executive, the person notes what they changed and why. It takes seconds, it gives a manager something specific to review, and it shows you who is handing back what the model produced without adding to it.
- Commit the recovered hours before the tool ships. Decide in advance what the freed-up time is for, more candidate conversations or deeper hiring manager calibration, and write it into the rollout plan next to the hours you expect to save. Unassigned hours go back into existing work, where the saving does not show up anywhere.
- Write AI-augmented work into the expectations for three roles. Pick the three roles where AI touches the most work and add one line to each describing what strong work with it looks like. Until that language sits in the document that decides raises, these behaviors stay optional.
Fluency with the tools is no longer what separates people. What separates them is a set of behaviors you can watch, teach, and write into a job expectation, once you have a live view of who is already doing them.