Saved in:
Bibliographic Details
Main Authors: Dagli, Rishit, Berger, Guillaume, Materzynska, Joanna, Bax, Ingo, Memisevic, Roland
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.02921
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917793909178368
author Dagli, Rishit
Berger, Guillaume
Materzynska, Joanna
Bax, Ingo
Memisevic, Roland
author_facet Dagli, Rishit
Berger, Guillaume
Materzynska, Joanna
Bax, Ingo
Memisevic, Roland
contents We introduce AirLetters, a new video dataset consisting of real-world videos of human-generated, articulated motions. Specifically, our dataset requires a vision model to predict letters that humans draw in the air. Unlike existing video datasets, accurate classification predictions for AirLetters rely critically on discerning motion patterns and on integrating long-range information in the video over time. An extensive evaluation of state-of-the-art image and video understanding models on AirLetters shows that these methods perform poorly and fall far behind a human baseline. Our work shows that, despite recent progress in end-to-end video understanding, accurate representations of complex articulated motions -- a task that is trivial for humans -- remains an open problem for end-to-end learning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02921
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AirLetters: An Open Video Dataset of Characters Drawn in the Air
Dagli, Rishit
Berger, Guillaume
Materzynska, Joanna
Bax, Ingo
Memisevic, Roland
Computer Vision and Pattern Recognition
We introduce AirLetters, a new video dataset consisting of real-world videos of human-generated, articulated motions. Specifically, our dataset requires a vision model to predict letters that humans draw in the air. Unlike existing video datasets, accurate classification predictions for AirLetters rely critically on discerning motion patterns and on integrating long-range information in the video over time. An extensive evaluation of state-of-the-art image and video understanding models on AirLetters shows that these methods perform poorly and fall far behind a human baseline. Our work shows that, despite recent progress in end-to-end video understanding, accurate representations of complex articulated motions -- a task that is trivial for humans -- remains an open problem for end-to-end learning.
title AirLetters: An Open Video Dataset of Characters Drawn in the Air
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.02921