Advancing Egocentric Video Question Answering with Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Patel, Alkesh, Chitalia, Vibhav, Yang, Yinfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
von: Park, Yohan, et al.
Veröffentlicht: (2025)
von: Park, Yohan, et al.
Veröffentlicht: (2025)
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
Towards Autonomous Instrument Tray Assembly for Sterile Processing Applications
von: Sankaranarayanan, Raghavasimhan, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Raghavasimhan, et al.
Veröffentlicht: (2026)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
Explore until Confident: Efficient Exploration for Embodied Question Answering
von: Ren, Allen Z., et al.
Veröffentlicht: (2024)
von: Ren, Allen Z., et al.
Veröffentlicht: (2024)
Whole-Body Conditioned Egocentric Video Prediction
von: Bai, Yutong, et al.
Veröffentlicht: (2025)
von: Bai, Yutong, et al.
Veröffentlicht: (2025)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
von: Wu, Tao, et al.
Veröffentlicht: (2025)
von: Wu, Tao, et al.
Veröffentlicht: (2025)
LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
von: Patel, Alkesh, et al.
Veröffentlicht: (2026)
von: Patel, Alkesh, et al.
Veröffentlicht: (2026)
Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models
von: Dong, Hao, et al.
Veröffentlicht: (2025)
von: Dong, Hao, et al.
Veröffentlicht: (2025)
EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
von: Srivastava, Varun, et al.
Veröffentlicht: (2025)
von: Srivastava, Varun, et al.
Veröffentlicht: (2025)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
3D Hand Pose Estimation in Everyday Egocentric Images
von: Prakash, Aditya, et al.
Veröffentlicht: (2023)
von: Prakash, Aditya, et al.
Veröffentlicht: (2023)
TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning
von: Spigler, Giacomo
Veröffentlicht: (2026)
von: Spigler, Giacomo
Veröffentlicht: (2026)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
von: Li, Qixiu, et al.
Veröffentlicht: (2025)
von: Li, Qixiu, et al.
Veröffentlicht: (2025)
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
von: Gao, Shenyuan, et al.
Veröffentlicht: (2026)
von: Gao, Shenyuan, et al.
Veröffentlicht: (2026)
Towards Fine-Grained Video Question Answering
von: Dai, Wei, et al.
Veröffentlicht: (2025)
von: Dai, Wei, et al.
Veröffentlicht: (2025)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
von: Wang, Beichen, et al.
Veröffentlicht: (2024)
von: Wang, Beichen, et al.
Veröffentlicht: (2024)
LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
World Simulation with Video Foundation Models for Physical AI
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
Know Where You're Uncertain When Planning with Multimodal Foundation Models: A Formal Framework
von: Bhatt, Neel P., et al.
Veröffentlicht: (2024)
von: Bhatt, Neel P., et al.
Veröffentlicht: (2024)
Turning Video Models into Generalist Robot Policies
von: Li, Sizhe Lester, et al.
Veröffentlicht: (2026)
von: Li, Sizhe Lester, et al.
Veröffentlicht: (2026)
OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation
von: Jawaid, Ahad, et al.
Veröffentlicht: (2025)
von: Jawaid, Ahad, et al.
Veröffentlicht: (2025)
Deep Dependency Networks and Advanced Inference Schemes for Multi-Label Classification
von: Arya, Shivvrat, et al.
Veröffentlicht: (2024)
von: Arya, Shivvrat, et al.
Veröffentlicht: (2024)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
von: Feng, Yao, et al.
Veröffentlicht: (2025)
von: Feng, Yao, et al.
Veröffentlicht: (2025)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
von: Min, Juhong, et al.
Veröffentlicht: (2024)
von: Min, Juhong, et al.
Veröffentlicht: (2024)
Advancing Autonomous Vehicle Intelligence: Deep Learning and Multimodal LLM for Traffic Sign Recognition and Robust Lane Detection
von: Sah, Chandan Kumar, et al.
Veröffentlicht: (2025)
von: Sah, Chandan Kumar, et al.
Veröffentlicht: (2025)
Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering
von: Frahm, Noah, et al.
Veröffentlicht: (2025)
von: Frahm, Noah, et al.
Veröffentlicht: (2025)
Grounding Video Models to Actions through Goal Conditioned Exploration
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
von: Li, Yingwei, et al.
Veröffentlicht: (2026)
von: Li, Yingwei, et al.
Veröffentlicht: (2026)
ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling
von: Zhou, Zikang, et al.
Veröffentlicht: (2024)
von: Zhou, Zikang, et al.
Veröffentlicht: (2024)
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
von: Gao, Haoxiang, et al.
Veröffentlicht: (2025)
von: Gao, Haoxiang, et al.
Veröffentlicht: (2025)
LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
von: Sha, Hao, et al.
Veröffentlicht: (2023)
von: Sha, Hao, et al.
Veröffentlicht: (2023)
Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition
von: Xi, Xinyu, et al.
Veröffentlicht: (2025)
von: Xi, Xinyu, et al.
Veröffentlicht: (2025)
Scaling Spatial Intelligence with Multimodal Foundation Models
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
AI Guide Dog: Egocentric Path Prediction on Smartphone
von: Jadhav, Aishwarya, et al.
Veröffentlicht: (2025)
von: Jadhav, Aishwarya, et al.
Veröffentlicht: (2025)
Visual Question Decomposition on Multimodal Large Language Models
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
von: Xiao, Junbin, et al.
Veröffentlicht: (2026) -
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
von: Yang, Ruihan, et al.
Veröffentlicht: (2025) -
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
von: Park, Yohan, et al.
Veröffentlicht: (2025) -
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
von: Wu, Tao, et al.
Veröffentlicht: (2024) -
Towards Autonomous Instrument Tray Assembly for Sterile Processing Applications
von: Sankaranarayanan, Raghavasimhan, et al.
Veröffentlicht: (2026)