EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Ye, Wong, Chi Kit, Lyu, Yuanhuiyi, Li, Hanqian, Huo, Jiahao, Chen, Jiacheng, Jiang, Lutao, Zheng, Xu, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
by: Peng, Taiying, et al.
Published: (2025)
by: Peng, Taiying, et al.
Published: (2025)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
by: Huo, Jiahao, et al.
Published: (2025)
by: Huo, Jiahao, et al.
Published: (2025)
BrightDreamer: Generic 3D Gaussian Generative Framework for Fast Text-to-3D Synthesis
by: Jiang, Lutao, et al.
Published: (2024)
by: Jiang, Lutao, et al.
Published: (2024)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
by: Zhu, Bingwen, et al.
Published: (2026)
by: Zhu, Bingwen, et al.
Published: (2026)
RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
by: Zhang, Deheng, et al.
Published: (2025)
by: Zhang, Deheng, et al.
Published: (2025)
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
by: Yan, Jiaqi, et al.
Published: (2025)
by: Yan, Jiaqi, et al.
Published: (2025)
EgoAVU: Egocentric Audio-Visual Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
by: Liu, Shaonan, et al.
Published: (2026)
by: Liu, Shaonan, et al.
Published: (2026)
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
by: Lyu, Yuanhuiyi, et al.
Published: (2026)
by: Lyu, Yuanhuiyi, et al.
Published: (2026)
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
by: Liao, Chenfei, et al.
Published: (2025)
by: Liao, Chenfei, et al.
Published: (2025)
EgoPrompt: Prompt Learning for Egocentric Action Recognition
by: Lyu, Huaihai, et al.
Published: (2025)
by: Lyu, Huaihai, et al.
Published: (2025)
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning
by: Wang, Zeyu, et al.
Published: (2026)
by: Wang, Zeyu, et al.
Published: (2026)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026)
by: Nagrani, Arsha, et al.
Published: (2026)
EgoVLM: Policy Optimization for Egocentric Video Understanding
by: Vinod, Ashwin, et al.
Published: (2025)
by: Vinod, Ashwin, et al.
Published: (2025)
EgoGen: An Egocentric Synthetic Data Generator
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
by: Plizzari, Chiara, et al.
Published: (2025)
by: Plizzari, Chiara, et al.
Published: (2025)
EgoSplat: Open-Vocabulary Egocentric Scene Understanding with Language Embedded 3D Gaussian Splatting
by: Li, Di, et al.
Published: (2025)
by: Li, Di, et al.
Published: (2025)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
by: Cheng, Sijie, et al.
Published: (2024)
by: Cheng, Sijie, et al.
Published: (2024)
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
by: Li, Yuan-Ming, et al.
Published: (2024)
by: Li, Yuan-Ming, et al.
Published: (2024)
Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
by: Wu, Tz-Ying, et al.
Published: (2024)
by: Wu, Tz-Ying, et al.
Published: (2024)
EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
by: Sun, Shitong, et al.
Published: (2026)
by: Sun, Shitong, et al.
Published: (2026)
EventBind: Learning a Unified Representation to Bind Them All for Event-based Open-world Understanding
by: Zhou, Jiazhou, et al.
Published: (2023)
by: Zhou, Jiazhou, et al.
Published: (2023)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
by: Wang, Ziyang, et al.
Published: (2026)
by: Wang, Ziyang, et al.
Published: (2026)
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
by: Leonardi, Rosario, et al.
Published: (2026)
by: Leonardi, Rosario, et al.
Published: (2026)
CompoNeRF: Text-guided Multi-object Compositional NeRF with Editable 3D Scene Layout
by: Bai, Haotian, et al.
Published: (2023)
by: Bai, Haotian, et al.
Published: (2023)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024)
by: Hummel, Thomas, et al.
Published: (2024)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
by: Ma, Jianzhe, et al.
Published: (2026)
by: Ma, Jianzhe, et al.
Published: (2026)
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
by: Liu, Yunqi, et al.
Published: (2026)
by: Liu, Yunqi, et al.
Published: (2026)
EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
by: Li, Jungang, et al.
Published: (2024)
by: Li, Jungang, et al.
Published: (2024)
EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment
by: Tang, Qiance, et al.
Published: (2026)
by: Tang, Qiance, et al.
Published: (2026)
Learning Modality-agnostic Representation for Semantic Segmentation from Any Modalities
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Image Anything: Towards Reasoning-coherent and Training-free Multi-modal Image Generation
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
Similar Items
-
Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
by: Lyu, Yuanhuiyi, et al.
Published: (2025) -
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
by: Zheng, Xu, et al.
Published: (2024) -
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
by: Zheng, Xu, et al.
Published: (2025) -
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
by: Peng, Taiying, et al.
Published: (2025) -
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
by: Huo, Jiahao, et al.
Published: (2025)