The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rossetto, Luca, Bailer, Werner, Dang-Nguyen, Duc-Tien, Healy, Graham, Jónsson, Björn Þór, Kongmeesub, Onanong, Le, Hoang-Bao, Rudinac, Stevan, Schöffmann, Klaus, Spiess, Florian, Tran, Allie, Tran, Minh-Triet, Tran, Quang-Linh, Gurrin, Cathal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908768058474496
author Rossetto, Luca
Bailer, Werner
Dang-Nguyen, Duc-Tien
Healy, Graham
Jónsson, Björn Þór
Kongmeesub, Onanong
Le, Hoang-Bao
Rudinac, Stevan
Schöffmann, Klaus
Spiess, Florian
Tran, Allie
Tran, Minh-Triet
Tran, Quang-Linh
Gurrin, Cathal
author_facet Rossetto, Luca
Bailer, Werner
Dang-Nguyen, Duc-Tien
Healy, Graham
Jónsson, Björn Þór
Kongmeesub, Onanong
Le, Hoang-Bao
Rudinac, Stevan
Schöffmann, Klaus
Spiess, Florian
Tran, Allie
Tran, Minh-Triet
Tran, Quang-Linh
Gurrin, Cathal
contents Egocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection containing ego- and exo-centric (i.e., first- and third-person perspective) video and audio from 15 time-aligned sources, as well as other sensor streams and auxiliary data. The dataset was recorded by volunteer participants over four days in a fixed location and includes the point of view of 10 participants, with an additional 5 fixed cameras providing an exocentric perspective. The entire dataset contains over 600 hours of UHD video recorded at 50 frames per second. In contrast to other datasets, CASTLE 2024 does not contain any partial censoring, such as blurred faces or distorted audio. The dataset is available via https://castle-dataset.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17116
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
Rossetto, Luca
Bailer, Werner
Dang-Nguyen, Duc-Tien
Healy, Graham
Jónsson, Björn Þór
Kongmeesub, Onanong
Le, Hoang-Bao
Rudinac, Stevan
Schöffmann, Klaus
Spiess, Florian
Tran, Allie
Tran, Minh-Triet
Tran, Quang-Linh
Gurrin, Cathal
Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
Egocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection containing ego- and exo-centric (i.e., first- and third-person perspective) video and audio from 15 time-aligned sources, as well as other sensor streams and auxiliary data. The dataset was recorded by volunteer participants over four days in a fixed location and includes the point of view of 10 participants, with an additional 5 fixed cameras providing an exocentric perspective. The entire dataset contains over 600 hours of UHD video recorded at 50 frames per second. In contrast to other datasets, CASTLE 2024 does not contain any partial censoring, such as blurred faces or distorted audio. The dataset is available via https://castle-dataset.github.io/.
title The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
topic Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
url https://arxiv.org/abs/2503.17116