HP Inc.

09/24/2026 | Press release | Distributed by Public on 09/24/2026 10:09

The Future of AI Is Hybrid, and the CIOs Who See It First Will Scale It Best

Almost every conversation with business and technology leaders I have these days finds its way to the same topic: AI tokenomics.

AI costs are already reshaping the infrastructure landscape. We are seeing growing adoption of lower-cost open-weight and open-source models among startups, developers and hobbyists, particularly for agentic workflows. At the same time, AI providers are evolving their pricing models as the economics of AI come into sharper focus.

As IT leaders, who spent the last two years scaling AI into every workflow now face rising cloud costs, they are asking a fair question: "Does all this AI really need to run in the cloud?"

The investment itself isn't the issue. Companies of every size are putting real money into AI, and our own research shows that 63% of workers now use AI in their workflows daily or weekly. The tools keep getting better. They're becoming more embedded into systems and processes and increasingly agentic. This is exactly what we hoped AI would become.

But once AI moves from a chat window into embedded workflows and autonomous agents, the requirements change. Cost is the complaint I hear first, and no wonder: CNBC reported an estimated 95% of enterprise AI workloads still rely on expensive frontier models, even though many routine tasks could be handled by smaller, lower-cost alternatives.

The second topic that comes up is security. Leaders are rethinking security for a world where autonomous agents, acting in real time across their systems, need answers instantly. Sensitive data raises hard questions about what should leave the building at all.

As AI capabilities scale, organizations face a critical tension: appetite for more intelligence, faster decisions and wider access to AI versus greater expense and more complex security risk.

This is driving demand for local inferencing that offers better economics, enhanced security, and lower latency.

The future of AI is hybrid

The good news? The industry is already moving in this direction. New silicon innovations are bringing meaningful AI compute directly onto the device, making it possible to enable AI to run more inference locally without every request going back to the cloud or adding another metered cost. My own HP Elitebook X is running a 20B model and routinely churning out useful information. This is an important shift: much of AI's next chapter will happen closer to the user, closer to the data and closer to where work actually gets done.

Computers once filled entire rooms; then they moved onto desks, into bags, into pockets - evolving at every step. AI is following the same path, only faster. Models are getting smaller and cheaper to run, as the devices already on desks and factory floors become powerful enough to run them.

Many AI workloads run better on AI PCs, workstations, and edge devices. This is especially true for agentic and inference workloads that need low latency, carry strict security requirements or draw on data that already lives at the edge. Running them locally delivers better token economics, lower latency, and stronger data security.

And there are additional benefits. When AI runs on the device, it creates experiences that are unique to that product. Context, memory and workflows stay close to the user, making the experience more personalized and secure. The value shifts from the cloud to the device itself.

This is not a retreat from the cloud. Cloud-based inference will remain essential for the largest, most complex, and most compute-intensive workloads. But the next phase of AI will require more intelligent orchestration, where each workload is matched to the environment that delivers the best balance of performance, economics, security, and user experience.

It's already happening

The real promise of AI at the edge is to enable people to do their best work, where it is needed: in clinics, classrooms, factories, field work, and everyday business workflows.

From NASA to the University of Kansas Health System, organizations are already running complex AI on the floor, using computer vision and audio models to identify defects in real time. The data is generated there, and the decision needs to happen there. Sending every frame to a data center and waiting for an answer makes no sense when milliseconds matter.

Analysts are seeing the same shift. Gartner projects more than 20% of enterprises will run AI workloads locally by 2028, and McKinsey expects inference to make up more than half of all AI workloads by 2030 with a growing share moving to the edge.

To make this real at enterprise scale, organizations will need more than powerful devices. They will need a complete hybrid AI stack: secure edge endpoints, model management, governance, AI security, and orchestration across devices, edge, and cloud. The question is no longer whether AI will move closer to where work happens. It is how quickly the industry can make that shift secure, scalable, and practical for customers. So, my advice to IT leaders working through this transition is to start with a different question - not where should the data be processed, but where does the work get done?

The best AI will be available where work happens - whether that is a factory line, an engineer's workstation or the laptop in front of you. Secure, personalized, on-device intelligence that reduces cloud cost while improving privacy, latency, and user productivity.

The future of AI is hybrid. The opportunity for enterprises is not to choose between cloud and edge, but to orchestrate intelligence where it creates the most value. That means designing AI architectures around the realities of work: where data is created, where decisions need to happen, and where employees and customers need better experiences. That is the future enterprises should be building toward now.

HP Inc. published this content on September 24, 2026, and is solely responsible for the information contained herein. Distributed via Public Technologies (PUBT), unedited and unaltered, on September 24, 2026 at 16:10 UTC. If you believe the information included in the content is inaccurate or outdated and requires editing or removal, please contact us at [email protected]