The first serious workplace studies of generative AI did not reveal one universal productivity multiplier. They revealed a slope.
At the lower end of the experience curve, an AI assistant could act like an always-available coach, supplying phrases, procedures and links at the moment a worker needed them. At the top, the same assistant often repeated knowledge the strongest employees already possessed. The result was not just higher output. It was a narrower gap between novices and veterans.
The finding that crystallised this pattern came from more than 5,000 customer-support agents. The widely circulated 2023 version reported a 34 per cent productivity increase among novice and lower-skilled workers, compared with minimal effects among experienced and highly skilled colleagues.
That number needs a version label. The study was later revised and peer-reviewed, changing some totals and estimates while preserving the central pattern. The careful conclusion is not that AI always gives beginners a 34 per cent boost. It is that, in one unusually large real-world deployment, the largest gains flowed to the people with the least experience and weakest baseline performance.
The famous 34 per cent figure came from the working-paper stage
Erik Brynjolfsson of Stanford and Danielle Li and Lindsey Raymond of MIT studied the staggered rollout of a generative AI assistant at a Fortune 500 company selling business-process software. Their 2023 NBER working paper covered 5,179 customer-support agents.
It reported that access to the assistant increased issues resolved per hour by 14 per cent on average. Among novice and lower-skilled workers, the increase was 34 per cent. The most experienced and highly skilled workers saw minimal gains.
The peer-reviewed version published in the Quarterly Journal of Economics in 2025 used data from 5,172 agents. It estimated a 15 per cent average increase in successful resolutions per hour and about a 30 per cent gain among less-skilled and less-experienced agents.
The change from 34 to roughly 30 per cent is not a contradiction. Papers develop as samples, specifications and analyses are refined. It does mean the headline figure should be attributed to the earlier version. The final paper’s message remained the same: the effect varied sharply by starting skill and experience.
The assistant did not replace the agent
The tool monitored text chats between agents and customers and generated possible responses. It could suggest a sequence of diagnostic questions, recommend language or surface a link to technical documentation. Agents remained responsible for the conversation and were free to accept, edit or ignore what appeared.
Productivity was measured as successfully resolved customer issues per hour. That combined three elements: how long an individual chat lasted, how many chats an agent handled and whether the customer’s problem was resolved. The 15 per cent result was not a subjective rating supplied by the AI vendor.
The rollout was staggered rather than a simple random assignment. Managers helped decide when teams and agents received training and access. The researchers used difference-in-differences models with controls for individual agents, calendar time and job tenure, then ran alternative estimators and an instrumental-variable analysis based on team rollout timing.
Those checks make the result more credible, but they do not turn it into a universal law. It was one tool inside one company, supporting one occupation with a relatively stable product and a large archive of previous chats.
AI compressed months of learning into weeks
The clearest picture came from the experience curves. Agents without the assistant began at roughly 1.8 successful resolutions per hour and took eight to ten months to reach about 2.5. Agents with AI from their first month reached the same level after around two months and continued improving.
Across several outcomes, two months of tenure with AI produced performance comparable to more than six months of tenure without it. The system did not erase learning. It steepened the early part of the curve.
The likely mechanism was knowledge transfer. The model had been trained on earlier customer conversations, including examples produced by high-performing agents. It could deliver those recurring strategies to a newcomer inside a live chat rather than waiting for a manager to select the conversation for a later coaching session.
Moderately uncommon problems produced the largest gains. For very common questions, even novices usually knew what to do. For extremely rare questions, the AI had too little training data. Between those extremes, the model had enough examples to help while the human agent often lacked firsthand experience.
The evidence suggests some learning survived the tool
A productivity assistant can create two very different outcomes. Workers may learn from repeated recommendations, or they may become dependent on a system that does the remembering for them.
The support study found suggestive evidence for learning. During occasional system outages, agents who had spent longer working with AI continued to handle chats faster than their own pre-AI baseline. The effect appeared to strengthen with the length of prior exposure.
That finding is not conclusive. Outages were rare, did not affect every worker equally and were not designed as clean experiments. Still, it suggests at least some of the model’s advice became human knowledge rather than disappearing when the suggestions stopped.
Customers also became more polite and less likely to ask for a manager after AI deployment. The paper associated access with lower employee turnover, especially among newer workers, although the authors were more cautious about the attrition analysis because each worker can leave only once and access was not randomly assigned.
Other early studies found the same tilt, with an important warning
The call-centre result was not isolated. In a preregistered experiment involving 453 college-educated professionals, Shakked Noy and Whitney Zhang randomly gave half the participants access to ChatGPT for occupation-specific writing tasks. Average completion time fell by 40 per cent and independently rated quality rose by 18 per cent. Workers with weaker initial skills benefited more, compressing the productivity distribution.
A separate experiment assigned 758 Boston Consulting Group consultants to work with or without GPT-4. On tasks inside the model’s competence, people below the median on a baseline assessment improved by 43 per cent, compared with 17 per cent for the upper half. Even in a highly selected professional workforce, the weaker performers gained more.
But the consulting study also demonstrated the “jagged technological frontier.” On a problem deliberately chosen to sit outside GPT-4’s strengths, consultants with AI were 19 per cent less likely to produce a correct answer. AI compressed the skill gap where it worked and punished misplaced trust where it did not.
A smaller controlled study of GitHub Copilot found that developers completed a JavaScript programming task 55.8 per cent faster with the tool. The heterogeneous results again suggested larger benefits for less-experienced developers, although a short coding exercise is far removed from maintaining a production system.
Compressing performance is not the same as eliminating expertise
Generative AI is especially good at redistributing codified patterns. If the answer has appeared often enough in training data and can be expressed as text, a model can place it in front of a beginner without years of repetition. That raises the performance floor.
Expertise becomes more visible at the edge. Experienced workers are better placed to recognise when the retrieved pattern does not fit, when a customer’s case is genuinely new, when the confident answer is wrong or when optimising speed damages quality. In the support study, the most skilled agents registered small declines in some measures of conversation quality while using AI.
There is also a circularity problem. High performers generated many of the examples that made the assistant useful. If they follow its adequate suggestions instead of developing better solutions, the future training pool may become less original. If firms treat shared AI output as proof that expert knowledge has no value, they may weaken the source of the next improvement.
For managers, this changes the purpose of senior roles rather than automatically removing them. Coaching, quality control, exception handling and creating new practice may become more important as routine expertise is distributed more widely.
A smaller output gap does not guarantee a smaller wage gap
The customer-support study did not measure wages, total employment or hiring decisions. A firm might respond to faster novice ramp-up by hiring more entry-level agents. It might reduce training costs, redesign jobs, raise service volume or decide it needs fewer people. The productivity result cannot tell us which path will dominate.
Silicon Canals previously examined how this support assistant narrowed the gap between new agents and veterans. The wider first wave of evidence points in the same direction when a task sits inside the model’s strengths: weaker workers often have more room to gain.
The boundary matters as much as the boost. AI can compress advantages built through repetition, retrieval and exposure to common cases. It does not automatically replace the judgement required when the pattern breaks. The new skill hierarchy may be flatter in routine execution and steeper at deciding when the machine should not be trusted.




