Back to All
Developer Blog

How AI models reach the Hexagon NPU: QAIRT, FastRPC, Stubs & Skels explained

During a recent conversation, my colleague Olivier Bloch (Director of Product Management for Qualcomm Technologies, Inc.) told me that if I’m asked the same question more than three times, I should write a blog about it. This is now my litmus test, and after a year of being knee-deep in hackathons, meetings with developers and companies, one question has come up at least three times per event.

How do my models actually reach the Qualcomm Hexagon NPU?

Well, it’s time to break it down because the confusing part usually isn’t running inference. It’s understanding what happens between your application calling an API and things actually running on the Hexagon NPU.

The Fundamentals: RPC

First stop, what is RPC? Understanding RPC will get you most of the way there. RPC stands for remote procedure call. At a high level, it’s a communication mechanism that allows a user to invoke a procedure/function in another process, address space, or processor as though it were local.

Okay, too vague.

Let’s get closer to what we use it for, if you’ve checked out the new family of dual-brain Arduino products, UNO Q/VENTUNO Q, RPC allows the MPU (Linux side) to communicate with the MCU (Zephyr side).

The two sides agree on an interface, which functions are exposed on either side, what arguments are accepted and what’s returned. The RPC framework handles getting the request to the other processor and bringing back the results.

Below is an example of RPC being utilized within the Arduino App Lab:

Sign up

Sign up for Developer monthly newsletter

Join thousands of developers around the globe who receive latest news and updates from our monthly curated newsletter.

Qualcomm-image


Now, how does this all apply to Qualcomm Technologies? Enter FastRPC.

Gateway to the Hexagon NPU: FastRPC

Now let’s get into what we all came for.

FastRPC is the framework developed by Qualcomm Technologies that allows remote function calls to move across the boundary between the application processor and Hexagon DSP. With respect to the QAIRT path, we’ll get there, it provides the communication infrastructure used to reach the Hexagon Tensor Processor (HTP). 

Qualcomm-image

Sounds familiar? Well, it should because that’s exactly the RPC concept that was covered in the fundamentals section, except instead of communicating between two computers, or two separate SoCs, we’re communicating between two different processing domains within the same SoC.

This is intentionally a high-level view of FastRPC. If you want to go down the FastRPC rabbit hole, I encourage you to go check the FastRPC repository that Qualcomm Technologies maintains. I’ve gone down this rabbit hole myself more times than I’d like to admit.

For our purposes, the important thing to remember is this. The stub represents the application-processor side of the remote call. FastRPC gets that call and transports the request across the processor/Hexagon (HTP) boundary, and the skeleton (skel) receives and dispatches it inside the Hexagon NPU.

In a generic FastRPC application, the interface is defined through an IDL file and the corresponding stub, and skel code are generated for you. Using QAIRT you typically don’t generate these pieces yourself. Qualcomm Technologies provides the versioned QNN HTP stub and skel libraries that implement this contract for the HTP backend.

Now, let’s get more into implementation details.

Mobile

The necessary OEM system image generally provides the platform-level FastRPC driver, services, firmware, and vendor libraries

For mobile, specifically /vendor/lib64 /vendor/lib subdirectories. Below is an example of the subdirectory and a couple highlighted libraries that are used in the FastRPC framework

Join Developer Discord

Come for support, stay for the community

Get support from experts, connect with like-minded developers, and access exclusive virtual events.

Qualcomm-image

Windows on Snapdragon

On Windows-based Compute platforms, NPU access is provided by the MCDM driver, accompanied by FastRPC service libraries. MCDM and FastRPC serve distinct roles, so MCDM should not be considered a FastRPC driver.

The MCDM artifacts can be found at C:\Windows\System32\DriverStore\FileRepository\qcnspmcdm<xxx>.inf_arm64_<xxxxxxxxx>

Qualcomm-image

Industrial IoT

Okay, so now let’s talk about IoT (Arduino counts too), because on these platforms it may be worth verifying that the FastRPC infrastructure is available before proceeding. A quick place to start is checking for the FastRPC device nodes, userspace libraries, and CDSP-related kernel messages.

```bash
ls -l /dev/fastrpc* 2>/dev/null
find /usr/lib /lib -iname ‘*fastrpc*’ 2>/dev/null
dmesg | grep -iE ‘fastrpc|cdsp’
```

If the expected FastRPC device nodes are missing, moving forward with an HTP deployment is pointless as installing additional userspace libraries won’t solve this kind of problem. The kernel driver, DSP firmware, device configuration, and other platform specific components are generally provided by the board’s BSP or system image, so you’ll need to check the documentation for the specific board or image. With that said, you more than likely will have everything that’s needed especially if you’re using QLI, or a board-supported Ubuntu/Debian image.

Putting it all together

Now, we need to get into QAIRT SDK because this is the part you’ll more than likely need to interact with the most.

First, what is QAIRT SDK? QAIRT stands for Qualcomm AI Runtime, and it provides tools APIs, libraries, and runtime backends for optimizing and running AI models across supported platforms from Qualcomm Technologies. Depending on the target and selected backend, QAIRT allows you to execute workloads using CPU, GPU, HTP, LPAI, or whatever new processor is developed in the future.

For compiling and running models, one of the important things you need to know is the Hexagon (HTP) architecture version of your target chipset.  The Hexagon version is important because it will determine which version of HTP components you’ll use, such as the v73 stub and skel.

Remember from the FastRPC section the stub represents the application process side of the remote procedure call (RPC) whereas the skel represents the DSP side of the RPC.

Your host operating system and architecture are equally important as this will determine which host side QAIRT libraries you’ll need.

To summarize, we have two compatibility domains. Your host operating system and architecture will determine which QAIRT libraries you use while the target’s HTP version determines which host and DSP-side binaries you’ll use. The final host-side stub must satisfy both.

Below is an example targeting aarh64-oe-linux-gcc11.2. Within QAIRT the stub files are located in lib/<system_architecture>

Qualcomm-image

The DSP-side skel libraries are located under the directory corresponding to your target’s HTP architecture version. The example below shows the runtime components provided for v73.

Qualcomm-image

Okay, now we know where FastRPC components are as well as the location of the Skel and Stub libraries that complete the contract. The next natural question is, how do I find what’s the Hexagon version of my particular chipset! I’m going show you a few ways from our different runtimes, the added benefit is that if you’re using any of these runtimes you can also use these paths to see if your device is supported.

ExecuTorch:

<chipset> = <SoC Model> #<Hexagon Version>

Qualcomm-image
Qualcomm-image
Qualcomm-image

https://workbench.aihub.qualcomm.com/devices

Above are the main components you need to get started, now we’ll get into something that always comes up, QAIRT SDK version. Once you’ve gone to Qualcomm Software Center where you download QAIRT from you’ll notice multiple versions.

Best practice is to compile and run models using the same QAIRT version. With that said, supported targets (HTP versions) may provide backward compatibility that allow artifacts produced with an older QAIRT release to run on a newer runtime. The reverse is not something you should depend on, a model produced by a newer SDK may require capabilities that an older runtime doesn’t have.

Now bear with me as I make this concrete. Say you compile a model on a host machine with QAIRT version 2.50.40.XXX but then attempt to run on a target with QAIRT version 2.37.1.XXX this may or may not work.

If you compile the model on a host machine with QAIRT version 2.37.1.XXX and run this on a target that has QAIRT version 2.50.40.XXX the older artifact will likely run on the newer target runtime, although you won’t benefit from optimizations and features that were introduced with the newer QAIRT version.

One last caveat is that for older Hexagon versions this backward compatibility may not be guaranteed.

I should digress really quickly and point out the expected workflow for compiling and running models on platforms from Qualcomm Technologies. We first have to keep in mind that we’re likely deploying models to memory constrained devices. At best, compiling directly on these devices is inefficient; at worst it may not be possible.

The typical workflow is to compile a model on a host machine, then after the model is compiled send it to the target device for deployment.

Story Time: I was preparing to show the world how to run Llama Stories, a 110M parameter model, on the Qualcomm Dragonwing QCS6490 chipset using ExecuTorch. The day before the event I had the idea that I’d just compile this model on target. I’m eagerly watching the output from HTOP as RAM is nearing the max, I come to the realization that I should have added swap. Well at this point, I’m committed and there is no turning back especially since it’s been compiling for 4 hours already.

Watching, sweating, watching, sweating, “VICTORY” I yell!

A whole 7 hours and the artifact was mine; the compiled model was ready for deployment, and the event went off without a hitch.

I finally got back home from the trip and as an experiment decided to compile the same model on my desktop, a whole two minutes later I had my compiled model.

Long story short, just because you can doesn’t mean you should.

One last thing I need to bring up, and this is with respect to IoT. Sometimes the QNN HTP skel that is packaged with QAIRT SDK can’t be found or loaded. The logs will tell you that the skel failed to open, followed by failed dlopen, initialization errors, or even later segfaults. This can happen because the DSP search path is incorrect, the required runtime components are missing, or the host libraries, skel, target firmware, and compiled model are incompatible.

 If you’re using a board supported Ubuntu image for IoT my suggestion would be to pull the necessary skel/stubs from Ubuntu PPA for platforms from Qualcomm Technologies.

```bash
sudo apt install -y software-properties-common
sudo add-apt-repository ppa:ubuntu-qcom-iot/qcom-ppa
sudo apt update
sudo apt install -y libqnn1 libqnn-dev qnn-tools
```

Once installed, inspect the package contents and update ADSP_LIBRARY_PATH to point to the directory containing the compatible DSP-side runtime libraries. 

```bash
dpkg -L libqnn1 libqnn-dev qnn-tools | grep -Ei ‘QnnHtp|Stub|Skel’
```

This may resolve skel loading type issues but keep the QAIRT and HTP version compatibility requirements we discussed earlier in mind.

If your BSP provides the working FastRPC device nodes but doesn’t provide the required FastRPC userspace libraries, you can build these components from the FastRPC repository. This will not create missing kernel driver, device node, DSP subsystem, or platform configurations

To do this you’ll first follow the instructions at https://github.com/qualcomm/fastrpc to install FastRPC. Depending on your target and system image, you may also need the corresponding DSP support binaries. These are tied to the exact SoC and DSP firmware revision, so follow the instructions in the Hexagon DSP binaries repository https://github.com/linux-msm/hexagon-dsp-binaries/tree/trunk.

The finale

And that’s really it! Hopefully by now you will understand how everything fits together, how FastRPC is the transportation mechanism for getting messages from the application processor side to the Hexagon NPU.

All in all, the flow is your application calls a framework, the QNN HTP backend loads the appropriate host-side libraries, the stub packages the request, FastRPC carries it across the processor boundary, and the skel receives it inside the Hexagon NPU so the Hexagon Tensor Processor (HTP) can execute your model.

Some key things to remember, the exact files change based on your host platform, HTP version, and QAIRT release, but the workflow remains essentially the same.

So, next time you see that a skel failed to open at least you’ll have a good idea of what this means, why this happens, and some clear ways to debug what may be going on, or should I say your agent will!

Opinions expressed in the content posted here are the personal opinions of the original authors, and do not necessarily reflect those of Qualcomm Incorporated or its subsidiaries ("Qualcomm"). The content is provided for informational purposes only and is not meant to be an endorsement or representation by Qualcomm or any other party. This site may also provide links or references to non-Qualcomm sites and resources. Qualcomm makes no representations, warranties, or other commitments whatsoever about any non-Qualcomm sites or third-party resources that may be referenced, accessible from, or linked to this site.

Qualcomm branded products are products of Qualcomm Technologies, Inc. and/or its subsidiaries.

About the Author
Derrick  Johnson
Derrick JohnsonStaff Engineer and Developer Advocate at Qualcomm Technologies, Inc

© Qualcomm Technologies, Inc. and/or its affiliated companies.

Snapdragon and Qualcomm branded products are products of Qualcomm Technologies, Inc. and/or its subsidiaries. Qualcomm patents are licensed by Qualcomm Incorporated.

Note: Certain services and materials may require you to accept additional terms and conditions before accessing or using those items.

References to "Qualcomm" may mean Qualcomm Incorporated, or subsidiaries or business units within the Qualcomm corporate structure, as applicable.

Qualcomm Incorporated includes our licensing business, QTL, and the vast majority of our patent portfolio. Qualcomm Technologies, Inc., a subsidiary of Qualcomm Incorporated, operates, along with its subsidiaries, substantially all of our engineering, research and development functions, and substantially all of our products and services businesses, including our QCT semiconductor business.

Materials that are as of a specific date, including but not limited to press releases, presentations, blog posts and webcasts, may have been superseded by subsequent events or disclosures.

Nothing in these materials is an offer to sell or license any of the services or materials referenced herein.