Camera Movement Classification in Historical Footage: A Comparative Study of Deep Video Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912652780896256 |
|---|---|
| author | Lin, Tingyu Dadras, Armin Kleber, Florian Sablatnig, Robert |
| author_facet | Lin, Tingyu Dadras, Armin Kleber, Florian Sablatnig, Robert |
| contents | Camera movement conveys spatial and narrative information essential for understanding video content. While recent camera movement classification (CMC) methods perform well on modern datasets, their generalization to historical footage remains unexplored. This paper presents the first systematic evaluation of deep video CMC models on archival film material. We summarize representative methods and datasets, highlighting differences in model design and label definitions. Five standard video classification models are assessed on the HISTORIAN dataset, which includes expert-annotated World War II footage. The best-performing model, Video Swin Transformer, achieves 80.25% accuracy, showing strong convergence despite limited training data. Our findings highlight the challenges and potential of adapting existing models to low-quality video and motivate future work combining diverse input modalities and temporal architectures. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_14713 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Camera Movement Classification in Historical Footage: A Comparative Study of Deep Video Models Lin, Tingyu Dadras, Armin Kleber, Florian Sablatnig, Robert Computer Vision and Pattern Recognition Artificial Intelligence Image and Video Processing Camera movement conveys spatial and narrative information essential for understanding video content. While recent camera movement classification (CMC) methods perform well on modern datasets, their generalization to historical footage remains unexplored. This paper presents the first systematic evaluation of deep video CMC models on archival film material. We summarize representative methods and datasets, highlighting differences in model design and label definitions. Five standard video classification models are assessed on the HISTORIAN dataset, which includes expert-annotated World War II footage. The best-performing model, Video Swin Transformer, achieves 80.25% accuracy, showing strong convergence despite limited training data. Our findings highlight the challenges and potential of adapting existing models to low-quality video and motivate future work combining diverse input modalities and temporal architectures. |
| title | Camera Movement Classification in Historical Footage: A Comparative Study of Deep Video Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Image and Video Processing |
| url | https://arxiv.org/abs/2510.14713 |