AirLetters dataset
AirLetters dataset
chip image

The AirLetters dataset is a benchmark for evaluating a model’s ability to classify articulated motions.

This large collection of over 161,000 video-label pairs of video clips, shows humans drawing letters and digits in the air, and is used to evaluate a model’s ability to classify articulated motions correctly. Unlike existing video datasets, AirLetters’ accurate classification predictions rely on discerning motion patterns and integrating information presented by the video over time (i.e., over many frames of video).

Our study revealed that while trivial for humans, accurate representations of complex articulated motions remain an open problem for end-to-end learning for models. For a detailed description of the dataset characteristics and research applications, see AirLetters: An Open Video Dataset of Characters Drawn in the Air.

Samples from AirLetters

The following are sample video frames showing the detection and classification of letters drawn by humans:

Figure 1: Images showing samples of articulated motion and their correct classifications.

Dataset Details

Total number of video-label pairs

161,652

Total classes

36

Video Format

.mp4

Max video width

640 pixels

The AirLetters dataset has over 161,000 of video-label pairs, covering:

  • 36 primary gesture classes (26 Latin letters and 10 digits)
  • Two contrasting classes: “doing nothing” and “doing other things”


The data is provided in the following directory structure:

  • /airletters
    • /videos: Video files containing the participants' motions
    • test.csv: Test data
    • train.csv: Training data
    • val.csv: Validation data


Each CSV contains the following columns:

  • id: Unique data ID
  • filename: Name of the video file.
  • label: Description of the drawn letter/behavior in the video
  • worker_id: ID of the individual who recorded the video
  • video_duration: Length of video (in seconds)

Dataset Collection Process
 

A custom platform integrated with crowd-sourcing providers was used to collect our dataset. This platform enabled us to recruit participants from diverse gender, geographical, and ethnic backgrounds and to provide the required instruction and recording functionality.
 

Participants were asked to record themselves performing all 36 gestures in front of their camera. Detailed visual and textual instructions were provided to ensure clear hand visibility, high video quality, and precise gesture execution. Supplementary example videos were provided to demonstrate correct gestures and address text instructions' limitations.
 

After reviewing the guidelines, participants prepared for recording with the help of a countdown timer. Recordings averaged around three seconds, after which participants could review and re-record if necessary. Each participant could make up to three submissions. To encourage scene variability, participants could interrupt recording and resume later.

Figure 2: Example of the diverse participants.

All submissions were reviewed by human operators to verify accuracy. Participants with mostly correct submissions but minor errors were allowed to make corrections and resubmit. This approach ensured the high quality and consistency of the dataset. Finally, all videos were resized to a width of 640 pixels, maintaining the aspect ratio.

Dataset license

 

The AirLetters dataset is available for research purposes.


Qualcomm AirLetters Dataset License Agreement - Research Use

Dataset Citation Instructions
 

If you use this dataset, please cite:

@inproceedings{dagliairletters,
    title={AirLetters: An Open Video Dataset of Characters Drawn in the Air},
    author={Dagli, Rishit and Berger, Guillaume and Materzynska, Joanna and Bax, Ingo and Memisevic, Roland},
    booktitle={European Conference on Computer Vision Workshops},
    year={2024}
}

Qualcomm AI Research
 

At Qualcomm AI Research, we are advancing AI to make its core capabilities – perception, reasoning, and action – ubiquitous across devices. Our mission is to make breakthroughs in fundamental AI research and scale them across industries. By bringing together some of the best minds in the field, we’re pushing the boundaries of what’s possible and shaping the future of AI.
 

Qualcomm AI Research continues to invest in and support deep-learning research in computer vision. The publication of the AirLetters dataset for use by the AI research community is one of our many initiatives.
 

Find out more about Qualcomm AI Research.
 

For any questions or technical support, please contact us at research.datasets@qti.qualcomm.com
 

Qualcomm AI Research is an initiative of Qualcomm Technologies, Inc.

Connect with our communities

Stay ahead of the curve

Receive the latest updates, exclusive offers, and valuable insights delivered through the Qualcomm newsletter straight to your inbox.

Stay ahead of the curve

Receive the latest updates, exclusive offers, and valuable insights delivered through the Qualcomm newsletter straight to your inbox.

© Qualcomm Technologies, Inc. and/or its affiliated companies.

Snapdragon and Qualcomm branded products are products of Qualcomm Technologies, Inc. and/or its subsidiaries. Qualcomm patented technologies are licensed by Qualcomm Incorporated.

Note: Certain services and materials may require you to accept additional terms and conditions before accessing or using those items.

References to "Qualcomm" may mean Qualcomm Incorporated, or subsidiaries or business units within the Qualcomm corporate structure, as applicable.

Qualcomm Incorporated includes our licensing business, QTL, and the vast majority of our patent portfolio. Qualcomm Technologies, Inc., a subsidiary of Qualcomm Incorporated, operates, along with its subsidiaries, substantially all of our engineering, research and development functions, and substantially all of our products and services businesses, including our QCT semiconductor business.

Materials that are as of a specific date, including but not limited to press releases, presentations, blog posts and webcasts, may have been superseded by subsequent events or disclosures.

Nothing in these materials is an offer to sell or license any of the services or materials referenced herein.