EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Deheng, Fu, Yuqian, Yang, Runyi, Miao, Yang, Qian, Tianwen, Zheng, Xu, Sun, Guolei, Chhatkuli, Ajad, Huang, Xuanjing, Jiang, Yu-Gang, Van Gool, Luc, Paudel, Danda Pani
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917304406638592
author Zhang, Deheng
Fu, Yuqian
Yang, Runyi
Miao, Yang
Qian, Tianwen
Zheng, Xu
Sun, Guolei
Chhatkuli, Ajad
Huang, Xuanjing
Jiang, Yu-Gang
Van Gool, Luc
Paudel, Danda Pani
author_facet Zhang, Deheng
Fu, Yuqian
Yang, Runyi
Miao, Yang
Qian, Tianwen
Zheng, Xu
Sun, Guolei
Chhatkuli, Ajad
Huang, Xuanjing
Jiang, Yu-Gang
Van Gool, Luc
Paudel, Danda Pani
contents Most existing benchmarks for understanding egocentric vision focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first comprehensive benchmark for nighttime egocentric vision, with visual question answering (VQA) as the core task. A key feature of EgoNight is the introduction of day-night aligned videos, which enhance night annotation quality using the daytime data and reveal clear performance gaps between lighting conditions. To achieve this, we collect both synthetic videos rendered by Blender and real-world recordings, ensuring that scenes and actions are visually and temporally aligned. Leveraging these paired videos, we construct EgoNight-VQA, supported by a novel day-augmented night auto-labeling engine and refinement through extensive human verification. Each QA pair is double-checked by annotators for reliability. In total, EgoNight-VQA contains 3658 QA pairs across 90 videos, spanning 12 diverse QA types, with more than 300 hours of human work. Evaluations of state-of-the-art multimodal large language models (MLLMs) reveal substantial performance drops when transferring from day to night, underscoring the challenges of reasoning under low-light conditions. Beyond VQA, EgoNight also introduces two auxiliary tasks, day-night correspondence retrieval and egocentric depth estimation at night, that further explore the boundaries of existing models. We believe EgoNight-VQA provides a strong foundation for advancing application-driven egocentric vision research and for developing models that generalize across illumination domains. The code and data can be found at https://github.com/dehezhang2/EgoNight.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06218
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
Zhang, Deheng
Fu, Yuqian
Yang, Runyi
Miao, Yang
Qian, Tianwen
Zheng, Xu
Sun, Guolei
Chhatkuli, Ajad
Huang, Xuanjing
Jiang, Yu-Gang
Van Gool, Luc
Paudel, Danda Pani
Computer Vision and Pattern Recognition
Artificial Intelligence
Most existing benchmarks for understanding egocentric vision focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first comprehensive benchmark for nighttime egocentric vision, with visual question answering (VQA) as the core task. A key feature of EgoNight is the introduction of day-night aligned videos, which enhance night annotation quality using the daytime data and reveal clear performance gaps between lighting conditions. To achieve this, we collect both synthetic videos rendered by Blender and real-world recordings, ensuring that scenes and actions are visually and temporally aligned. Leveraging these paired videos, we construct EgoNight-VQA, supported by a novel day-augmented night auto-labeling engine and refinement through extensive human verification. Each QA pair is double-checked by annotators for reliability. In total, EgoNight-VQA contains 3658 QA pairs across 90 videos, spanning 12 diverse QA types, with more than 300 hours of human work. Evaluations of state-of-the-art multimodal large language models (MLLMs) reveal substantial performance drops when transferring from day to night, underscoring the challenges of reasoning under low-light conditions. Beyond VQA, EgoNight also introduces two auxiliary tasks, day-night correspondence retrieval and egocentric depth estimation at night, that further explore the boundaries of existing models. We believe EgoNight-VQA provides a strong foundation for advancing application-driven egocentric vision research and for developing models that generalize across illumination domains. The code and data can be found at https://github.com/dehezhang2/EgoNight.
title EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.06218