EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Yuping, Huang, Yifei, Chen, Guo, Pei, Baoqi, Xu, Jilan, Lu, Tong, Pang, Jiangmiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
von: Xu, Jilan, et al.
Veröffentlicht: (2025)
von: Xu, Jilan, et al.
Veröffentlicht: (2025)
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
von: Huang, Yifei, et al.
Veröffentlicht: (2024)
von: Huang, Yifei, et al.
Veröffentlicht: (2024)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
von: He, Yuping, et al.
Veröffentlicht: (2025)
von: He, Yuping, et al.
Veröffentlicht: (2025)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
von: Jung, Minjoon, et al.
Veröffentlicht: (2025)
von: Jung, Minjoon, et al.
Veröffentlicht: (2025)
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
von: Grauman, Kristen, et al.
Veröffentlicht: (2023)
von: Grauman, Kristen, et al.
Veröffentlicht: (2023)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
EgoExo-WM: Unlocking Exo Video for Ego World Models
von: Tran, Danny, et al.
Veröffentlicht: (2026)
von: Tran, Danny, et al.
Veröffentlicht: (2026)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
von: Lin, Jingli, et al.
Veröffentlicht: (2025)
von: Lin, Jingli, et al.
Veröffentlicht: (2025)
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
Intention-driven Ego-to-Exo Video Generation
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
von: Zhang, Zixin, et al.
Veröffentlicht: (2025)
von: Zhang, Zixin, et al.
Veröffentlicht: (2025)
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
von: Park, Jungin, et al.
Veröffentlicht: (2025)
von: Park, Jungin, et al.
Veröffentlicht: (2025)
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
von: Liu, Ruiping, et al.
Veröffentlicht: (2026)
von: Liu, Ruiping, et al.
Veröffentlicht: (2026)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
von: Fu, Yuqian, et al.
Veröffentlicht: (2024)
von: Fu, Yuqian, et al.
Veröffentlicht: (2024)
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
von: Özsoy, Ege, et al.
Veröffentlicht: (2025)
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2024)
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2024)
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
von: Dai, Yang, et al.
Veröffentlicht: (2026)
von: Dai, Yang, et al.
Veröffentlicht: (2026)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
von: Wang, Yi, et al.
Veröffentlicht: (2024)
von: Wang, Yi, et al.
Veröffentlicht: (2024)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
von: Xu, Pengju, et al.
Veröffentlicht: (2025)
von: Xu, Pengju, et al.
Veröffentlicht: (2025)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
ENIGMA-360: An Ego-Exo Dataset for Human Behavior Understanding in Industrial Scenarios
von: Ragusa, Francesco, et al.
Veröffentlicht: (2026)
von: Ragusa, Francesco, et al.
Veröffentlicht: (2026)
SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
PCIE_EgoHandPose Solution for EgoExo4D Hand Pose Challenge
von: Chen, Feng, et al.
Veröffentlicht: (2024)
von: Chen, Feng, et al.
Veröffentlicht: (2024)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
von: Zhang, Hongjie, et al.
Veröffentlicht: (2023)
von: Zhang, Hongjie, et al.
Veröffentlicht: (2023)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
von: Lin, Jingli, et al.
Veröffentlicht: (2025)
von: Lin, Jingli, et al.
Veröffentlicht: (2025)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
von: Lin, Junming, et al.
Veröffentlicht: (2024)
von: Lin, Junming, et al.
Veröffentlicht: (2024)
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
von: Li, Caorui, et al.
Veröffentlicht: (2025)
von: Li, Caorui, et al.
Veröffentlicht: (2025)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
Retrieval-Augmented Egocentric Video Captioning
von: Xu, Jilan, et al.
Veröffentlicht: (2024)
von: Xu, Jilan, et al.
Veröffentlicht: (2024)
UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs
von: Jiang, Lifan, et al.
Veröffentlicht: (2026)
von: Jiang, Lifan, et al.
Veröffentlicht: (2026)
Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
von: Mahdi, Mohammad, et al.
Veröffentlicht: (2025)
von: Mahdi, Mohammad, et al.
Veröffentlicht: (2025)
Robust Ego-Exo Correspondence with Long-Term Memory
von: Hu, Yijun, et al.
Veröffentlicht: (2025)
von: Hu, Yijun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
von: Xu, Jilan, et al.
Veröffentlicht: (2025) -
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
von: Pei, Baoqi, et al.
Veröffentlicht: (2025) -
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024) -
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
von: Huang, Yifei, et al.
Veröffentlicht: (2024) -
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)