EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Boshen, Mei, Yuting, Liu, Xinbi, Zheng, Sipeng, Jin, Qin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View World
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
SPAFormer: Sequential 3D Part Assembly with Transformers
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation
von: Zhang, Libo, et al.
Veröffentlicht: (2025)
von: Zhang, Libo, et al.
Veröffentlicht: (2025)
EgoVLM: Policy Optimization for Egocentric Video Understanding
von: Vinod, Ashwin, et al.
Veröffentlicht: (2025)
von: Vinod, Ashwin, et al.
Veröffentlicht: (2025)
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2024)
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2024)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
EgoLM: Multi-Modal Language Model of Egocentric Motions
von: Hong, Fangzhou, et al.
Veröffentlicht: (2024)
von: Hong, Fangzhou, et al.
Veröffentlicht: (2024)
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
von: Zhang, Deheng, et al.
Veröffentlicht: (2025)
von: Zhang, Deheng, et al.
Veröffentlicht: (2025)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
von: Valdez, Hector A., et al.
Veröffentlicht: (2024)
von: Valdez, Hector A., et al.
Veröffentlicht: (2024)
EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
von: Fu, Hongming, et al.
Veröffentlicht: (2026)
von: Fu, Hongming, et al.
Veröffentlicht: (2026)
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
von: Mei, Yuting, et al.
Veröffentlicht: (2024)
von: Mei, Yuting, et al.
Veröffentlicht: (2024)
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
von: Tian, Shulin, et al.
Veröffentlicht: (2025)
von: Tian, Shulin, et al.
Veröffentlicht: (2025)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
EgoSelf: From Memory to Personalized Egocentric Assistant
von: Wang, Yanshuo, et al.
Veröffentlicht: (2026)
von: Wang, Yanshuo, et al.
Veröffentlicht: (2026)
EgoGen: An Egocentric Synthetic Data Generator
von: Li, Gen, et al.
Veröffentlicht: (2024)
von: Li, Gen, et al.
Veröffentlicht: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
von: Li, Runjia, et al.
Veröffentlicht: (2025)
von: Li, Runjia, et al.
Veröffentlicht: (2025)
EgoM2P: Egocentric Multimodal Multitask Pretraining
von: Li, Gen, et al.
Veröffentlicht: (2025)
von: Li, Gen, et al.
Veröffentlicht: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
von: Cheng, Sijie, et al.
Veröffentlicht: (2024)
von: Cheng, Sijie, et al.
Veröffentlicht: (2024)
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
von: Cheng, Qinchuan, et al.
Veröffentlicht: (2026)
von: Cheng, Qinchuan, et al.
Veröffentlicht: (2026)
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
von: Kim, Kangsan, et al.
Veröffentlicht: (2026)
von: Kim, Kangsan, et al.
Veröffentlicht: (2026)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
von: Bhalgat, Yash, et al.
Veröffentlicht: (2024)
von: Bhalgat, Yash, et al.
Veröffentlicht: (2024)
EMAG: Ego-motion Aware and Generalizable 2D Hand Forecasting from Egocentric Videos
von: Hatano, Masashi, et al.
Veröffentlicht: (2024)
von: Hatano, Masashi, et al.
Veröffentlicht: (2024)
EgoLife: Towards Egocentric Life Assistant
von: Yang, Jingkang, et al.
Veröffentlicht: (2025)
von: Yang, Jingkang, et al.
Veröffentlicht: (2025)
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms
von: VanVoorst, Brian, et al.
Veröffentlicht: (2026)
von: VanVoorst, Brian, et al.
Veröffentlicht: (2026)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
von: Lall, Vishakha, et al.
Veröffentlicht: (2025)
von: Lall, Vishakha, et al.
Veröffentlicht: (2025)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video
von: Tan, Yuting, et al.
Veröffentlicht: (2026)
von: Tan, Yuting, et al.
Veröffentlicht: (2026)
X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
von: Wang, Yunzhe, et al.
Veröffentlicht: (2025)
TrajMamba: An Ego-Motion-Guided Mamba Model for Pedestrian Trajectory Prediction from an Egocentric Perspective
von: Peng, Yusheng, et al.
Veröffentlicht: (2026)
von: Peng, Yusheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View World
von: Xu, Boshen, et al.
Veröffentlicht: (2024) -
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
von: Xu, Boshen, et al.
Veröffentlicht: (2024) -
SPAFormer: Sequential 3D Part Assembly with Transformers
von: Xu, Boshen, et al.
Veröffentlicht: (2024) -
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026) -
EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation
von: Zhang, Libo, et al.
Veröffentlicht: (2025)