The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908768058474496 |
|---|---|
| author | Rossetto, Luca Bailer, Werner Dang-Nguyen, Duc-Tien Healy, Graham Jónsson, Björn Þór Kongmeesub, Onanong Le, Hoang-Bao Rudinac, Stevan Schöffmann, Klaus Spiess, Florian Tran, Allie Tran, Minh-Triet Tran, Quang-Linh Gurrin, Cathal |
| author_facet | Rossetto, Luca Bailer, Werner Dang-Nguyen, Duc-Tien Healy, Graham Jónsson, Björn Þór Kongmeesub, Onanong Le, Hoang-Bao Rudinac, Stevan Schöffmann, Klaus Spiess, Florian Tran, Allie Tran, Minh-Triet Tran, Quang-Linh Gurrin, Cathal |
| contents | Egocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection containing ego- and exo-centric (i.e., first- and third-person perspective) video and audio from 15 time-aligned sources, as well as other sensor streams and auxiliary data. The dataset was recorded by volunteer participants over four days in a fixed location and includes the point of view of 10 participants, with an additional 5 fixed cameras providing an exocentric perspective. The entire dataset contains over 600 hours of UHD video recorded at 50 frames per second. In contrast to other datasets, CASTLE 2024 does not contain any partial censoring, such as blurred faces or distorted audio. The dataset is available via https://castle-dataset.github.io/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_17116 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding Rossetto, Luca Bailer, Werner Dang-Nguyen, Duc-Tien Healy, Graham Jónsson, Björn Þór Kongmeesub, Onanong Le, Hoang-Bao Rudinac, Stevan Schöffmann, Klaus Spiess, Florian Tran, Allie Tran, Minh-Triet Tran, Quang-Linh Gurrin, Cathal Multimedia Artificial Intelligence Computer Vision and Pattern Recognition Information Retrieval Egocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection containing ego- and exo-centric (i.e., first- and third-person perspective) video and audio from 15 time-aligned sources, as well as other sensor streams and auxiliary data. The dataset was recorded by volunteer participants over four days in a fixed location and includes the point of view of 10 participants, with an additional 5 fixed cameras providing an exocentric perspective. The entire dataset contains over 600 hours of UHD video recorded at 50 frames per second. In contrast to other datasets, CASTLE 2024 does not contain any partial censoring, such as blurred faces or distorted audio. The dataset is available via https://castle-dataset.github.io/. |
| title | The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding |
| topic | Multimedia Artificial Intelligence Computer Vision and Pattern Recognition Information Retrieval |
| url | https://arxiv.org/abs/2503.17116 |