Back to All
OnQ Blog

Why agentic AI needs a completely different mobile architecture: the Qualcomm Hexagon NPU

Qualcomm-image



What you should know:
  • Agentic AI demands a fundamentally different mobile architecture — one built around specialized models routed by task and context rather than a single massive model — and the foundation for it has to be designed in from the start.
  • Our new Qualcomm Hexagon NPU is built for that shift, led by two moves that matter — a transformer-focused Element Accelerator and a 50% larger shared memory that keeps frequently accessed model data close to the NPU — so agents stay responsive as they juggle longer context, more tools and concurrent tasks.
  • We designed it to run that experience on device: Mixture-of-Experts models that activate a fraction of their parameters per token, a broad INT2–FP16 precision range, and up to 50% faster prefill for INT4 models — bringing responsive, private generative AI to the phone.



The next era of mobile computing will be defined by agents: intelligent systems that understand personal context, operate across applications, reason continuously and take action on your behalf.

As AI becomes more agentic, the winning architecture will not be one massive model handling every task. It will be a system of specialized models, routed intelligently based on the task, the context and the user.

That shift demands more than faster AI inference. It requires a new architecture built for persistent, multimodal, low-latency intelligence. With our next-generation premium mobile platform, Qualcomm Technologies is introducing that foundation, led by the new Qualcomm Hexagon NPU and two standout features: the new Element Accelerator and a much larger shared memory system.

The Element Accelerator is purpose-built for the transformer workloads that power modern generative and agentic AI. Together with the scalar, vector and matrix extensions, it accelerates the operations that matter most for large models, helping agents respond faster, reason more efficiently and deliver richer experiences without compromising mobile power efficiency.

Just as important, our Hexagon NPU significantly expands its large shared memory. A 50% larger NPU memory subsystem allows more model state, activations and intermediate tensors to remain on-chip, minimizing trips to DDR. This reduces memory bottlenecks and delivers faster, more responsive agentic AI experiences.

"With our next-generation premium mobile platform, Qualcomm Technologies is introducing that foundation, led by the new Qualcomm Hexagon NPU and two standout features: the new Element Accelerator and a much larger shared memory system."

Together, these innovations define the role of the next-generation Hexagon NPU: running more of the agentic AI experience directly on device. The Element Accelerator, vector, matrix and scalar extensions, and larger shared memory work as a collective architecture, helping the NPU accelerate transformer workloads and bring blazing fast agentic experiences to the device.

 

What the Hexagon NPU delivers

For agents, performance is about keeping intelligence responsive, contextual and efficient across multiple tasks. The Hexagon NPU is designed for always-running AI, long-context reasoning, multimodal models, concurrent agents and low-latency action loops.

Expanded NPU memory capacity allows more model data, activations and intermediate tensors to remain local to the processor, reducing trips to external memory and helping agents answer, revise, plan and act with less waiting. That means richer context windows, faster token generation and a smoother experience as agents move between tools and tasks.

The Element Accelerator further improves performance on the transformer operations that matter most for generative and agentic AI, while vector extensions help drive high-throughput AI math and scalar extensions to support the decision logic, routing and orchestration that agents increasingly require.

 

Running larger models with less memory bandwidth

The Hexagon NPU is also built for the next generation of model architectures. Qualcomm Technologies is partnering with leading memory and model providers to incorporate a new generation of on-device AI architectures built around Mixture-of-Experts (MoE) models. A 30B-parameter MoE can keep tens of billions of parameters available while activating only about 3B routed parameters for each token generation step on the NPU.

Unlike dense models, which require a much larger portion of the network to participate in every inference step, MoE dynamically selects specialized experts based on the input, reducing active compute and memory bandwidth demands. Combined with intelligent flash-to-memory expert management and caching techniques, this approach enables larger-model-class AI experiences with low power consumption and memory requirements, bringing responsive, private generative AI to phones. MoE models selectively activate specialized expert networks, improving efficiency and quality without invoking an entire model for every request.

"The Hexagon NPU is also built for the next generation of model architectures."

That approach complements the broader precision strategy. Support for INT2, INT4, INT8, FP8 and FP16 gives developers more flexibility to balance performance, memory, model quality and power. Smaller, more efficient models can handle more local tasks, while larger models can be accessed when the workload truly requires them.

For INT4-based models, the platform delivers up to 50% higher prefill performance, faster decoding throughput, enhanced speculative decoding and higher overall tokens per second. The result is a more immediate on-device AI experience, with agents that feel faster, more capable and more personal.

 

Built to run agentic AI on device

With the next-generation Hexagon NPU, our latest premium mobile processor is designed to make agentic AI more local, efficient and responsive. It brings together acceleration, memory, precision and model-loading innovations so more intelligence can run close to the user.

At the center of it all is the Hexagon NPU: a collective architecture with the Element Accelerator, vector and scalar extensions, larger shared memory and flash-enabled model loading built to deliver the next generation of on-device AI.

The Agentic Era has arrived. Our next-generation premium mobile processor is built to run it on device.




Go Deeper
How do the CPU and NPU work together to run agentic AI on-device?

Our Hexagon NPU is built to accelerate transformer workloads and keep context close to compute, but agents also need intelligent routing, orchestration and data availability as tasks move between cores. The CPU plays a complementary role in orchestrating those multi-step workloads and keeping data ready for the NPU. 

Does AI acceleration extend beyond the NPU to the rest of the platform?

It does — the architectural principles behind our Hexagon NPU, where specialized hardware and local memory work together to improve performance, latency and power efficiency, are becoming a platform-wide capability. AI processing is expanding into the GPU as well, with dedicated acceleration that mirrors the NPU's approach. 

Where can I learn more about how these technologies fit into Qualcomm Technologies' broader vision?

This post focuses on the Hexagon NPU, but the Element Accelerator, our larger shared memory and support for MoE models are part of a larger next-generation mobile platform story. 

Opinions expressed in the content posted here are the personal opinions of the original authors, and do not necessarily reflect those of Qualcomm Incorporated or its subsidiaries ("Qualcomm"). The content is provided for informational purposes only and is not meant to be an endorsement or representation by Qualcomm or any other party. This site may also provide links or references to non-Qualcomm sites and resources. Qualcomm makes no representations, warranties, or other commitments whatsoever about any non-Qualcomm sites or third-party resources that may be referenced, accessible from, or linked to this site.

Qualcomm branded products are products of Qualcomm Technologies, Inc. and/or its subsidiaries.

About the Author
Vinesh Sukumar
Vinesh SukumarVP, Product Management of AI/GenAI, Qualcomm Technologies, Inc.

© Qualcomm Technologies, Inc. and/or its affiliated companies.

Snapdragon and Qualcomm branded products are products of Qualcomm Technologies, Inc. and/or its subsidiaries. Qualcomm patents are licensed by Qualcomm Incorporated.

Note: Certain services and materials may require you to accept additional terms and conditions before accessing or using those items.

References to "Qualcomm" may mean Qualcomm Incorporated, or subsidiaries or business units within the Qualcomm corporate structure, as applicable.

Qualcomm Incorporated includes our licensing business, QTL, and the vast majority of our patent portfolio. Qualcomm Technologies, Inc., a subsidiary of Qualcomm Incorporated, operates, along with its subsidiaries, substantially all of our engineering, research and development functions, and substantially all of our products and services businesses, including our QCT semiconductor business.

Materials that are as of a specific date, including but not limited to press releases, presentations, blog posts and webcasts, may have been superseded by subsequent events or disclosures.

Nothing in these materials is an offer to sell or license any of the services or materials referenced herein.