Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.02921 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917793909178368 |
|---|---|
| author | Dagli, Rishit Berger, Guillaume Materzynska, Joanna Bax, Ingo Memisevic, Roland |
| author_facet | Dagli, Rishit Berger, Guillaume Materzynska, Joanna Bax, Ingo Memisevic, Roland |
| contents | We introduce AirLetters, a new video dataset consisting of real-world videos of human-generated, articulated motions. Specifically, our dataset requires a vision model to predict letters that humans draw in the air. Unlike existing video datasets, accurate classification predictions for AirLetters rely critically on discerning motion patterns and on integrating long-range information in the video over time. An extensive evaluation of state-of-the-art image and video understanding models on AirLetters shows that these methods perform poorly and fall far behind a human baseline. Our work shows that, despite recent progress in end-to-end video understanding, accurate representations of complex articulated motions -- a task that is trivial for humans -- remains an open problem for end-to-end learning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_02921 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | AirLetters: An Open Video Dataset of Characters Drawn in the Air Dagli, Rishit Berger, Guillaume Materzynska, Joanna Bax, Ingo Memisevic, Roland Computer Vision and Pattern Recognition We introduce AirLetters, a new video dataset consisting of real-world videos of human-generated, articulated motions. Specifically, our dataset requires a vision model to predict letters that humans draw in the air. Unlike existing video datasets, accurate classification predictions for AirLetters rely critically on discerning motion patterns and on integrating long-range information in the video over time. An extensive evaluation of state-of-the-art image and video understanding models on AirLetters shows that these methods perform poorly and fall far behind a human baseline. Our work shows that, despite recent progress in end-to-end video understanding, accurate representations of complex articulated motions -- a task that is trivial for humans -- remains an open problem for end-to-end learning. |
| title | AirLetters: An Open Video Dataset of Characters Drawn in the Air |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.02921 |