Visual Intention Grounding for Egocentric Assistants
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Pengzhan, Xiao, Junbin, Tse, Tze Ho Elden, Li, Yicong, Akula, Arjun, Yao, Angela |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Human Motion Plausibility with Body Momentum
von: Nguyen, Ha Linh, et al.
Veröffentlicht: (2025)
von: Nguyen, Ha Linh, et al.
Veröffentlicht: (2025)
DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction
von: Xu, Kai, et al.
Veröffentlicht: (2024)
von: Xu, Kai, et al.
Veröffentlicht: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
A Constrained Optimization Approach for Gaussian Splatting from Coarsely-posed Images and Noisy Lidar Point Clouds
von: Peng, Jizong, et al.
Veröffentlicht: (2025)
von: Peng, Jizong, et al.
Veröffentlicht: (2025)
Humans as Checkerboards: Calibrating Camera Motion Scale for World-Coordinate Human Mesh Recovery
von: Yang, Fengyuan, et al.
Veröffentlicht: (2024)
von: Yang, Fengyuan, et al.
Veröffentlicht: (2024)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
Leveraging RGB Images for Pre-Training of Event-Based Hand Pose Estimation
von: Liu, Ruicong, et al.
Veröffentlicht: (2025)
von: Liu, Ruicong, et al.
Veröffentlicht: (2025)
TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction
von: Huang, Yiyao, et al.
Veröffentlicht: (2025)
von: Huang, Yiyao, et al.
Veröffentlicht: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
SA-GS: Semantic-Aware Gaussian Splatting for Large Scene Reconstruction with Geometry Constrain
von: Xiong, Butian, et al.
Veröffentlicht: (2024)
von: Xiong, Butian, et al.
Veröffentlicht: (2024)
Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics
von: Tse, Tze Ho Elden, et al.
Veröffentlicht: (2025)
von: Tse, Tze Ho Elden, et al.
Veröffentlicht: (2025)
GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose Refinement
von: Zheng, Linfang, et al.
Veröffentlicht: (2024)
von: Zheng, Linfang, et al.
Veröffentlicht: (2024)
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
von: Cheng, Qinchuan, et al.
Veröffentlicht: (2026)
von: Cheng, Qinchuan, et al.
Veröffentlicht: (2026)
High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation
von: Feng, Runyang, et al.
Veröffentlicht: (2025)
von: Feng, Runyang, et al.
Veröffentlicht: (2025)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
von: Yan, Jiaqi, et al.
Veröffentlicht: (2025)
von: Yan, Jiaqi, et al.
Veröffentlicht: (2025)
Analyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose Estimation
von: Zhao, Zhuoran, et al.
Veröffentlicht: (2025)
von: Zhao, Zhuoran, et al.
Veröffentlicht: (2025)
EgoLife: Towards Egocentric Life Assistant
von: Yang, Jingkang, et al.
Veröffentlicht: (2025)
von: Yang, Jingkang, et al.
Veröffentlicht: (2025)
Question-Answering Dense Video Events
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
On the Consistency of Video Large Language Models in Temporal Comprehension
von: Jung, Minjoon, et al.
Veröffentlicht: (2024)
von: Jung, Minjoon, et al.
Veröffentlicht: (2024)
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
von: Li, Lei-lei, et al.
Veröffentlicht: (2025)
von: Li, Lei-lei, et al.
Veröffentlicht: (2025)
Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
von: Li, Junlong, et al.
Veröffentlicht: (2025)
von: Li, Junlong, et al.
Veröffentlicht: (2025)
VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation
von: Liao, Xinyao, et al.
Veröffentlicht: (2026)
von: Liao, Xinyao, et al.
Veröffentlicht: (2026)
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
Intention-Conditioned Long-Term Human Egocentric Action Forecasting
von: Mascaro, Esteve Valls, et al.
Veröffentlicht: (2022)
von: Mascaro, Esteve Valls, et al.
Veröffentlicht: (2022)
Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs
von: Chen, Tianle, et al.
Veröffentlicht: (2025)
von: Chen, Tianle, et al.
Veröffentlicht: (2025)
EgoSelf: From Memory to Personalized Egocentric Assistant
von: Wang, Yanshuo, et al.
Veröffentlicht: (2026)
von: Wang, Yanshuo, et al.
Veröffentlicht: (2026)
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2024)
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2024)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
Grounded Question-Answering in Long Egocentric Videos
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2023)
von: Kundu, Sanjoy, et al.
Veröffentlicht: (2023)
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
von: Huang, Yifei, et al.
Veröffentlicht: (2024)
von: Huang, Yifei, et al.
Veröffentlicht: (2024)
HRVDA: High-Resolution Visual Document Assistant
von: Liu, Chaohu, et al.
Veröffentlicht: (2024)
von: Liu, Chaohu, et al.
Veröffentlicht: (2024)
WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos
von: Ye, Yufei, et al.
Veröffentlicht: (2026)
von: Ye, Yufei, et al.
Veröffentlicht: (2026)
VideoQA in the Era of LLMs: An Empirical Study
von: Xiao, Junbin, et al.
Veröffentlicht: (2024)
von: Xiao, Junbin, et al.
Veröffentlicht: (2024)
Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
von: Li, Ling, et al.
Veröffentlicht: (2026)
von: Li, Ling, et al.
Veröffentlicht: (2026)
Fine-grained Spatiotemporal Grounding on Egocentric Videos
von: Liang, Shuo, et al.
Veröffentlicht: (2025)
von: Liang, Shuo, et al.
Veröffentlicht: (2025)
REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
von: He, Qiyuan, et al.
Veröffentlicht: (2025)
von: He, Qiyuan, et al.
Veröffentlicht: (2025)
Simultaneous Detection and Interaction Reasoning for Object-Centric Action Recognition
von: Li, Xunsong, et al.
Veröffentlicht: (2024)
von: Li, Xunsong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improving Human Motion Plausibility with Body Momentum
von: Nguyen, Ha Linh, et al.
Veröffentlicht: (2025) -
DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction
von: Xu, Kai, et al.
Veröffentlicht: (2024) -
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023) -
A Constrained Optimization Approach for Gaussian Splatting from Coarsely-posed Images and Noisy Lidar Point Clouds
von: Peng, Jizong, et al.
Veröffentlicht: (2025) -
Humans as Checkerboards: Calibrating Camera Motion Scale for World-Coordinate Human Mesh Recovery
von: Yang, Fengyuan, et al.
Veröffentlicht: (2024)