EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jung, Minjoon, Xiao, Junbin, Kim, Junghyun, Zhang, Byoung-Tak, Yao, Angela
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914124006424576
author Jung, Minjoon
Xiao, Junbin
Kim, Junghyun
Zhang, Byoung-Tak
Yao, Angela
author_facet Jung, Minjoon
Xiao, Junbin
Kim, Junghyun
Zhang, Byoung-Tak
Yao, Angela
contents Can Video-LLMs achieve consistent temporal understanding when videos capture the same event from different viewpoints? To study this, we introduce EgoExo-Con (Consistency), a benchmark of comprehensively synchronized egocentric and exocentric video pairs with human-refined queries in natural language. EgoExo-Con emphasizes two temporal understanding tasks: Temporal Verification and Temporal Grounding. It evaluates not only correctness but consistency across viewpoints. Our analysis reveals two critical limitations of existing Video-LLMs: (1) models often fail to maintain consistency, with results far worse than their single-view performances. (2) When naively finetuned with synchronized videos of both viewpoints, the models show improved consistency but often underperform those trained on a single view. For improvements, we propose View-GRPO, a novel reinforcement learning framework that effectively strengthens view-specific temporal reasoning while encouraging consistent comprehension across viewpoints. Our method demonstrates its superiority over naive SFT and GRPO, especially for improving cross-view consistency. All resources will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26113
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
Jung, Minjoon
Xiao, Junbin
Kim, Junghyun
Zhang, Byoung-Tak
Yao, Angela
Computer Vision and Pattern Recognition
Artificial Intelligence
Can Video-LLMs achieve consistent temporal understanding when videos capture the same event from different viewpoints? To study this, we introduce EgoExo-Con (Consistency), a benchmark of comprehensively synchronized egocentric and exocentric video pairs with human-refined queries in natural language. EgoExo-Con emphasizes two temporal understanding tasks: Temporal Verification and Temporal Grounding. It evaluates not only correctness but consistency across viewpoints. Our analysis reveals two critical limitations of existing Video-LLMs: (1) models often fail to maintain consistency, with results far worse than their single-view performances. (2) When naively finetuned with synchronized videos of both viewpoints, the models show improved consistency but often underperform those trained on a single view. For improvements, we propose View-GRPO, a novel reinforcement learning framework that effectively strengthens view-specific temporal reasoning while encouraging consistent comprehension across viewpoints. Our method demonstrates its superiority over naive SFT and GRPO, especially for improving cross-view consistency. All resources will be made publicly available.
title EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.26113