Reading Recognition in the Wild

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Charig, Alam, Samiul, Siam, Shakhrul Iman, Proulx, Michael J., Mathias, Lambert, Somasundaram, Kiran, Pesqueira, Luis, Fort, James, Sheriffdeen, Sheroze, Parkhi, Omkar, Ren, Carl, Zhang, Mi, Chai, Yuning, Newcombe, Richard, Kim, Hyo Jin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915968855310336
author Yang, Charig
Alam, Samiul
Siam, Shakhrul Iman
Proulx, Michael J.
Mathias, Lambert
Somasundaram, Kiran
Pesqueira, Luis
Fort, James
Sheriffdeen, Sheroze
Parkhi, Omkar
Ren, Carl
Zhang, Mi
Chai, Yuning
Newcombe, Richard
Kim, Hyo Jin
author_facet Yang, Charig
Alam, Samiul
Siam, Shakhrul Iman
Proulx, Michael J.
Mathias, Lambert
Somasundaram, Kiran
Pesqueira, Luis
Fort, James
Sheriffdeen, Sheroze
Parkhi, Omkar
Ren, Carl
Zhang, Mi
Chai, Yuning
Newcombe, Richard
Kim, Hyo Jin
contents To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24848
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reading Recognition in the Wild
Yang, Charig
Alam, Samiul
Siam, Shakhrul Iman
Proulx, Michael J.
Mathias, Lambert
Somasundaram, Kiran
Pesqueira, Luis
Fort, James
Sheriffdeen, Sheroze
Parkhi, Omkar
Ren, Carl
Zhang, Mi
Chai, Yuning
Newcombe, Richard
Kim, Hyo Jin
Computer Vision and Pattern Recognition
Machine Learning
To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism.
title Reading Recognition in the Wild
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2505.24848