ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dang, Ronghao, Yuan, Yuqian, Zhang, Wenqi, Xin, Yifei, Zhang, Boqiang, Li, Long, Wang, Liuyi, Zeng, Qinyang, Li, Xin, Bing, Lidong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025)
RynnEC: Bringing MLLMs into Embodied World
von: Dang, Ronghao, et al.
Veröffentlicht: (2025)
von: Dang, Ronghao, et al.
Veröffentlicht: (2025)
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
von: Zhang, Boqiang, et al.
Veröffentlicht: (2025)
von: Zhang, Boqiang, et al.
Veröffentlicht: (2025)
Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents
von: Li, Long, et al.
Veröffentlicht: (2024)
von: Li, Long, et al.
Veröffentlicht: (2024)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
von: Yuan, Yuqian, et al.
Veröffentlicht: (2024)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2024)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
RynnBrain: Open Embodied Foundation Models
von: Dang, Ronghao, et al.
Veröffentlicht: (2026)
von: Dang, Ronghao, et al.
Veröffentlicht: (2026)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
von: Mo, Ying, et al.
Veröffentlicht: (2026)
von: Mo, Ying, et al.
Veröffentlicht: (2026)
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
von: Wang, Yifei, et al.
Veröffentlicht: (2025)
von: Wang, Yifei, et al.
Veröffentlicht: (2025)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
von: Suglia, Alessandro, et al.
Veröffentlicht: (2024)
von: Suglia, Alessandro, et al.
Veröffentlicht: (2024)
VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding
von: Liu, Zhaowei, et al.
Veröffentlicht: (2025)
von: Liu, Zhaowei, et al.
Veröffentlicht: (2025)
One-shot Robust Federated Learning of Independent Component Analysis
von: Jin, Dian, et al.
Veröffentlicht: (2025)
von: Jin, Dian, et al.
Veröffentlicht: (2025)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
Optimal vintage factor analysis with deflation varimax
von: Bing, Xin, et al.
Veröffentlicht: (2023)
von: Bing, Xin, et al.
Veröffentlicht: (2023)
From Perfect to Noisy World Simulation: Customizable Embodied Multi-modal Perturbations for SLAM Robustness Benchmarking
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning
von: Zhang, Min, et al.
Veröffentlicht: (2024)
von: Zhang, Min, et al.
Veröffentlicht: (2024)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
von: Cheng, Sijie, et al.
Veröffentlicht: (2024)
von: Cheng, Sijie, et al.
Veröffentlicht: (2024)
A Comprehensive Survey on World Models for Embodied AI
von: Li, Xinqing, et al.
Veröffentlicht: (2025)
von: Li, Xinqing, et al.
Veröffentlicht: (2025)
Exploring Chinese Music Students' AI Familiarity and AI Self‐Efficacy: A Mixed‐Methods Study Based on Social Cognitive Theory ( SCT )
von: Qinyang Sun
Veröffentlicht: (2026)
von: Qinyang Sun
Veröffentlicht: (2026)
Causality-based Cross-Modal Representation Learning for Vision-and-Language Navigation
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)
Vision-and-Language Navigation via Causal Learning
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)
von: Wang, Liuyi, et al.
Veröffentlicht: (2024)
WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
von: Shang, Yu, et al.
Veröffentlicht: (2026)
von: Shang, Yu, et al.
Veröffentlicht: (2026)
ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models
von: Peng, Baorui, et al.
Veröffentlicht: (2026)
von: Peng, Baorui, et al.
Veröffentlicht: (2026)
Beyond Turing: Memory-Amortized Inference as a Foundation for Cognitive Computation
von: Li, Xin
Veröffentlicht: (2025)
von: Li, Xin
Veröffentlicht: (2025)
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
von: Zhang, Deheng, et al.
Veröffentlicht: (2025)
von: Zhang, Deheng, et al.
Veröffentlicht: (2025)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
von: Zhou, Wenqi, et al.
Veröffentlicht: (2025)
von: Zhou, Wenqi, et al.
Veröffentlicht: (2025)
Nurses' Experiences and Attitudes Toward the Use of Nursing Robots: A Meta‐Synthesis of Qualitative Researches
von: Pan Yang, et al.
Veröffentlicht: (2025)
von: Pan Yang, et al.
Veröffentlicht: (2025)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models
von: Li, Zhong-Zhi, et al.
Veröffentlicht: (2024)
von: Li, Zhong-Zhi, et al.
Veröffentlicht: (2024)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
Local well-posedness of strong solutions to the compressible Navier-Stokes equations with degenerate viscosities and far field vacuum in 3D exterior domains
von: Li, Jiaxu, et al.
Veröffentlicht: (2026)
von: Li, Jiaxu, et al.
Veröffentlicht: (2026)
Strong solutions to the initial-boundary-value problem of compressible MHD equations with degenerate viscosities and far field vacuum in 3D exterior domains
von: Li, Jiaxu, et al.
Veröffentlicht: (2026)
von: Li, Jiaxu, et al.
Veröffentlicht: (2026)
WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
von: Shang, Yu, et al.
Veröffentlicht: (2026)
von: Shang, Yu, et al.
Veröffentlicht: (2026)
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025)
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025)
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025) -
RynnEC: Bringing MLLMs into Embodied World
von: Dang, Ronghao, et al.
Veröffentlicht: (2025) -
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
von: Zhang, Boqiang, et al.
Veröffentlicht: (2025) -
Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents
von: Li, Long, et al.
Veröffentlicht: (2024) -
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
von: Yuan, Yuqian, et al.
Veröffentlicht: (2024)