Build and ship On-Device AI
Deploy agentic and generative AI models directly with Snapdragon® and Qualcomm Dragonwing™ processors. Optimize for inference latency, cost, and power efficiency.
Artificial intelligence at your fingertips
Build with your preferred framework, deploy with Qualcomm
Develop models in TensorFlow, PyTorch, ONNX, or LiteRT and run them with hardware-accelerated inference across over four billion devices powered by Snapdragon and Dragonwing processors for mobile, compute, and IoT.
Choose your developer approach
Start with the Guided approach for pre-built workflows and framework runtimes or choose the Native approach for direct control over model execution, optimization, and runtime targeting.
Guided
Get started quickly with guided workflows and runtimes for ML frameworks.
Native
Native tools provide more direct control over optimization, runtime behavior, and model execution.
AI developer workflow
Sample Datasets
Sample datasets for training models across research, prototyping, and experimentation.
Datasets - Hugging Face
Access datasets for your use case on Hugging Face.
Download a pre-optimized model or bring your own and prepare it for on-device execution.
GUIDED
Pre-optimized models from Qualcomm AI Hub
Hundreds of pre-optimized ML models, Small Language Models (SLMs), and Large Language Models (LLMs) are ready to profile, test, and deploy.
Access models for your use case on Hugging Face.
NATIVE
Qualcomm AI Runtime SDK (QAIRT)
Prepare models from your own training data using Qualcomm® AI Runtime SDK (QAIRT). QAIRT includes conversion and optimization tools for TensorFlow, PyTorch, and ONNX, with execution across CPU, GPU, NPU, and Qualcomm® Sensing Hub on Snapdragon and Dragonwing.
Related Resource
Building and Executing Your Model tutorial which includes steps to build a model from data.
Qualcomm Neural Processing SDK
If you want to prepare a model from your own training data, start with the Qualcomm Neural Processing SDK which includes tools to build a compatible model for execution on Snapdragon or Dragonwing.
Choose from an array of devices with Snapdragon and Dragonwing for your compute, IoT, or mobile projects.
GUIDED
Access real compute and mobile devices with Dragonwing and Snapdragon remotely through your browser. Build, test, debug, and profile on physical devices without requiring physical devices.
Use Qualcomm AI Hub to profile and test on devices with Qualcomm AI Hub Workbench.
NATIVE
Windows on Snapdragon / Compute
Snapdragon X Series processors power the fastest, most power-efficient AI PCs from several major brands.
Explore the full range of Android OS devices powered by Snapdragon processors for your mobile apps and games.
Choose from powerful evaluation kit (EVK) options for your next agentic AI or GenAI project.
Convert your models into Snapdragon and Dragonwing compatible formats and optimize them for size, accuracy, and performance during execution.
GUIDED
Use Qualcomm AI Hub Workbench to quickly optimize a trained model for execution on over 60 cloud-based devices using supported runtimes.
Use Edge Impulse to build IoT datasets, train models, and optimize libraries to run directly on devices with Dragonwing.
NATIVE
AI Model Efficiency Toolkit (AIMET)
Use AIMET to quantize and compress models trained with PyTorch or ONNX before running them with ONNX Runtime (ORT), Qualcomm Neural Processing SDK, or Qualcomm AI Engine Direct SDK (QNN).
Qualcomm Neural Processing SDK includes several analysis and optimization tools to prepare a model for execution using the SDK’s runtime.
Related Resource
Quantization code samples introductory tutorial
Qualcomm Neural Processing SDK (QAIRT)
Qualcomm AI Runtime SDK (QAIRT) includes several analysis and optimization tools to prepare a model for execution using the SDK’s runtime.
Below are several runtimes you can use for hardware-accelerated inference in your application running on Snapdragon or Dragonwing.
GUIDED
Choose one of the many sample applications as your starting point for loading a model and running inference. See the Get Started with Qualcomm AI Hub Apps tutorial
Built on the Qualcomm AI Engine Direct SDK, these runtime integrations let you use your preferred AI/ML and Generative AI framework.
ONNX Runtime (ORT)
For ONNX models, ONNX Runtime (ORT) includes a Qualcomm Plugin Execution Provider that you can incorporate into your application.
LiteRT
To run LiteRT models incorporate the LiteRT Delegate into your application.
ExecuTorch
For PyTorch models, incorporate ExecuTorch into your application that includes a Qualcomm® Hexagon™ Delegate.
NATIVE
Qualcomm AI Runtime SDK (QAIRT)
Qualcomm AI Runtime SDK (QAIRT) is a suite of tools and runtimes to build, optimize, and run ML models generated by TensorFlow, PyTorch, and ONNX.
Related Resource
Building and Executing Your Model tutorial which includes steps to build a model from data.
Deploy your app, your model, or both and choose the tools that match your pipeline.
NATIVE
Qualcomm AI Runtime SDK (QAIRT)
Qualcomm AI Runtime SDK (QAIRT) is a suite of tools and runtimes to build, optimize, and run ML models generated by TensorFlow, PyTorch, and ONNX.
GUIDED
Choose one of the many sample applications as your starting point for loading a model and running inference. See the Get Started with Qualcomm AI Hub Apps tutorial
Use FoundriesFactory to build a CI/CD pipeline that manages OTA updates for apps and models to IoT edge devices in the field via containerization.
Use Cases
IoT
- Vision language models to assess remote areas
- Object analysis for healthcare or retail
- Predictive maintenance in factories
- Precision agriculture to monitor crops and soil conditions
Robotics
- Factory robots on assembly lines
- Voice-controlled robotic arms
- Autonomous restaurant servers
- Autonomous drones
- Automated healthcare assistants
Compute
- Fast photo and video editing
- Turn agentic IT ticket overload into a daily insight pipeline
- Voice-controlled user interfaces
- Multi-device agentic orchestration
Mobile
- Chatbots for documents and images
- Turn medical records into simple conversations
- Agentic web search, deep research and automation workflows
- On-demand translation and image generation
Gaming
- Intelligent and conversational non-playable characters (NPCs)
- Agent-based teammates
- In-game assistants
- Dynamic graphics or stories
XR
- Context-aware personal assistants
- Spatial perception and scene understanding
- Digital twins for enterprise and industrial environments
- Immersive training and education for teams
Connect with our communities
Stay ahead of the curve
Receive the latest updates, exclusive offers, and valuable insights delivered through the Qualcomm newsletter straight to your inbox.
Stay ahead of the curve
Receive the latest updates, exclusive offers, and valuable insights delivered through the Qualcomm newsletter straight to your inbox.
