09/11/2026 | Press release | Distributed by Public on 09/11/2026 11:52
Consulting firms spent the past two years encouraging employees to use artificial intelligence across their daily work. Now, as AI consumption reaches enormous levels, they are confronting a less visible challenge: how to control the cost of using the technology without undermining its productivity gains.
McKinsey is responding by giving employees greater visibility into how much AI they consume rather than imposing blanket limits on usage.
The consulting firm tracks AI consumption at the individual-user level and sends email alerts when an employee's usage becomes unusually high, according to Debasish Patnaik, who leads QuantumBlack, McKinsey's AI, data and analytics group in the UK.
Register for the next Tekedia Mini-MBA.
Register for Tekedia AI in Business Masterclass.
Join Tekedia Capital Syndicate and co-invest in great global startups.
Register for Nigeria Capital Market Masterclass.
The alert system was introduced across the firm during the summer.
"Similar to using mobile data on a work phone, we tell them this is how you could do things to make it more cost-effective for the firm," Patnaik said.
The shift points to a broader change taking place among large companies. Early enterprise AI strategies largely focused on getting employees to experiment with generative AI and demonstrate where the technology could improve productivity. As adoption has expanded, companies are discovering that widespread use can generate a substantial and recurring infrastructure bill.
The economics are becoming more complicated as AI providers increasingly charge for consumption. Instead of paying a simple flat subscription for access, companies can incur costs based on the number of tokens an AI model processes. Tokens are the small units of text that models read and generate. That means an employee who makes thousands of AI requests, uses lengthy context, or repeatedly asks a model to process large amounts of information can generate significantly more costs than another employee performing simpler tasks.
OpenAI said in September that its most prolific users of AI coding agents were consuming more than $7,000 worth of tokens a day, illustrating how quickly costs can rise when advanced AI tools are used intensively.
McKinsey's own consumption demonstrates the scale involved.
By May 2026, the firm was processing about five trillion AI tokens a month, according to a company blog post. Usage was highly concentrated, with roughly 10% of users accounting for about 65% of total consumption. Consultants and software engineers were among the heaviest users.
Rather than interpreting those figures as a reason to restrict access, McKinsey is using them to educate employees about the economics of AI.
"We really believe in giving autonomy to the consultants," Patnaik said.
The objective is to encourage employees to find more efficient ways of obtaining the same result. An employee might use a smaller model for a relatively simple task, reduce unnecessary context, or avoid repeatedly sending the same information to a model. That approach is based on a simple premise: the cost of AI should be evaluated alongside the value it creates rather than treated as an expense that must be minimized regardless of the outcome.
McKinsey says its internal AI spending has not yet reached a problematic level. Much of the firm's usage is aimed at improving individual productivity or delivering client engagements more quickly, cases where Patnaik said the benefits still outweigh the costs.
But that calculation could change as usage continues to increase.
"Usage could be more 'egregious' in another six months," Patnaik said, while noting that McKinsey has already introduced additional controls beyond individual usage alerts.
One is an internal AI gateway that optimizes requests before they reach external model providers. The firm has also introduced circuit breakers that can temporarily suspend access when token consumption becomes particularly high, allowing the company to determine whether the usage is generating sufficient value.
Caching provides another way to reduce expenditure. Responses to repeated questions can be reused rather than generated again, while consumption costs can be pooled across the business instead of tying unused capacity to individual licenses.
The measures point to a broader evolution in corporate AI management. Companies are beginning to treat AI consumption more like a variable operating expense that requires monitoring, optimization, and governance.
Other consulting firms are taking similar steps.
EY has established an "AI Value Realization Office" to oversee AI spending and has deployed an "invisible" routing system behind some specialized AI tools. The system directs requests to the model considered most appropriate for a particular task.
EY told Business Insider that the routing system, combined with other governance measures, had reduced token consumption by 60% since April.
At Deloitte, the economics of AI coding tools have also become an issue. A senior software engineer at Deloitte US told Business Insider in June that changes to GitHub's pricing model were "already wreaking havoc" on expectations for work, with developers quickly exhausting new monthly usage quotas.
The experience of consulting firms offers an early indication of a problem likely to spread across corporate America as AI moves deeper into business operations.
The initial enterprise AI question was whether companies could persuade employees to use the technology. The next question is whether companies can make widespread usage economically sustainable. The questions matter because AI costs do not necessarily rise in line with headcount. A relatively small group of heavy users can generate a disproportionate share of consumption, while increasingly capable models can also require more computing resources.
However, the situation creates a tension between controlling expenditure and preserving productivity for executives. Excessive restrictions could discourage employees from using AI for valuable tasks, while unrestricted usage could allow computational costs to grow faster than the benefits generated.
McKinsey is now taking its experience to clients.
Patnaik's primary role at QuantumBlack focuses on delivering AI capabilities to clients rather than managing McKinsey's internal AI use, but the firm established a formal practice this spring to advise companies on deploying AI more cost-effectively.
He said clients have increasingly recognized over the past quarter that there is a "hidden cost that we haven't completely thought through."
Companies therefore need to assess the competitive benefits of AI against the returns generated by that spending.
Patnaik argues that businesses should measure the cost of AI per outcome rather than simply calculating expenditure per employee. That distinction matters because reducing token consumption is not necessarily an improvement if the resulting AI system produces weaker work, takes employees longer to complete a task, or creates additional costs elsewhere.
"What you don't want is to take costs out here, but incur costs on the other side without knowing about it," he said.
The implication is that enterprise AI is entering a more mature phase. The priority is shifting from maximizing adoption to optimizing the relationship between AI consumption and measurable business outcomes.
For consulting firms that helped popularize corporate AI adoption, the transition is visible. The technology is no longer an experiment sitting on the edge of the organization. At companies such as McKinsey, it is already processing trillions of tokens every month.
The emerging challenge is making sure those tokens translate into enough additional productivity, revenue, or client value to justify the bill.