Saved in:
| Main Authors: | Li, Yifan, Zhou, Xinyu, Ge, Yunhao, Kong, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.20085 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Medical Visual Grounding via Knowledge-guided Spatial Prompts
by: Gao, Yifan, et al.
Published: (2026)
by: Gao, Yifan, et al.
Published: (2026)
Attention to Trajectory: Trajectory-Aware Open-Vocabulary Tracking
by: Li, Yunhao, et al.
Published: (2025)
by: Li, Yunhao, et al.
Published: (2025)
EgoNav: Egocentric Scene-aware Human Trajectory Prediction
by: Wang, Weizhuo, et al.
Published: (2024)
by: Wang, Weizhuo, et al.
Published: (2024)
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
by: Yoshida, Tomoya, et al.
Published: (2025)
by: Yoshida, Tomoya, et al.
Published: (2025)
Egocentric Event-Based Vision for Ping Pong Ball Trajectory Prediction
by: Alberico, Ivan, et al.
Published: (2025)
by: Alberico, Ivan, et al.
Published: (2025)
DreamDistribution: Learning Prompt Distribution for Diverse In-distribution Generation
by: Zhao, Brian Nlong, et al.
Published: (2023)
by: Zhao, Brian Nlong, et al.
Published: (2023)
I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners
by: Ling, Lu, et al.
Published: (2025)
by: Ling, Lu, et al.
Published: (2025)
Spatial-Conditioned Reasoning in Long-Egocentric Videos
by: Tribble, James, et al.
Published: (2026)
by: Tribble, James, et al.
Published: (2026)
EgoPrompt: Prompt Learning for Egocentric Action Recognition
by: Lyu, Huaihai, et al.
Published: (2025)
by: Lyu, Huaihai, et al.
Published: (2025)
MADiff: Motion-Aware Mamba Diffusion Models for Hand Trajectory Prediction on Egocentric Videos
by: Ma, Junyi, et al.
Published: (2024)
by: Ma, Junyi, et al.
Published: (2024)
HEADS-UP: Head-Mounted Egocentric Dataset for Trajectory Prediction in Blind Assistance Systems
by: Haghighi, Yasaman, et al.
Published: (2024)
by: Haghighi, Yasaman, et al.
Published: (2024)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
Visual Intention Grounding for Egocentric Assistants
by: Sun, Pengzhan, et al.
Published: (2025)
by: Sun, Pengzhan, et al.
Published: (2025)
EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
by: Bai, Yu, et al.
Published: (2026)
by: Bai, Yu, et al.
Published: (2026)
How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction
by: Jun, Sejoon, et al.
Published: (2026)
by: Jun, Sejoon, et al.
Published: (2026)
Intention Enhanced Diffusion Model for Multimodal Pedestrian Trajectory Prediction
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Attention to the Burstiness in Visual Prompt Tuning!
by: Wang, Yuzhu, et al.
Published: (2025)
by: Wang, Yuzhu, et al.
Published: (2025)
VPN: Visual Prompt Navigation
by: Feng, Shuo, et al.
Published: (2025)
by: Feng, Shuo, et al.
Published: (2025)
STF: Spatial Temporal Fusion for Trajectory Prediction
by: Han, Pengqian, et al.
Published: (2023)
by: Han, Pengqian, et al.
Published: (2023)
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos
by: Chen, Mingfei, et al.
Published: (2025)
by: Chen, Mingfei, et al.
Published: (2025)
Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
by: Zeng, Xianchao, et al.
Published: (2025)
by: Zeng, Xianchao, et al.
Published: (2025)
EgoKit: Towards Unified Low-Cost Egocentric Data Collection with Heterogeneous Devices
by: Yu, Liuchuan, et al.
Published: (2026)
by: Yu, Liuchuan, et al.
Published: (2026)
Retrieval-Enhanced Visual Prompt Learning for Few-shot Classification
by: Rong, Jintao, et al.
Published: (2023)
by: Rong, Jintao, et al.
Published: (2023)
Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation
by: Ge, Yunhao, et al.
Published: (2024)
by: Ge, Yunhao, et al.
Published: (2024)
EgoAVU: Egocentric Audio-Visual Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
Intention-Aware Diffusion Model for Pedestrian Trajectory Prediction
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
3D Skew Gaussian Splatting with Any Camera Trajectory Visualization Engine
by: Zhao, Beizhen, et al.
Published: (2026)
by: Zhao, Beizhen, et al.
Published: (2026)
VRP-SAM: SAM with Visual Reference Prompt
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning
by: Park, Sungjune, et al.
Published: (2026)
by: Park, Sungjune, et al.
Published: (2026)
Visual Attention Prompted Prediction and Learning
by: Zhang, Yifei, et al.
Published: (2023)
by: Zhang, Yifei, et al.
Published: (2023)
TP-DRSeg: Improving Diabetic Retinopathy Lesion Segmentation with Explicit Text-Prompts Assisted SAM
by: Li, Wenxue, et al.
Published: (2024)
by: Li, Wenxue, et al.
Published: (2024)
Visual Trajectory Prediction of Vessels for Inland Navigation
by: Puzicha, Alexander, et al.
Published: (2025)
by: Puzicha, Alexander, et al.
Published: (2025)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
by: Yun, Heeseung, et al.
Published: (2024)
by: Yun, Heeseung, et al.
Published: (2024)
UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Enhancing Prompt Following with Visual Control Through Training-Free Mask-Guided Diffusion
by: Chen, Hongyu, et al.
Published: (2024)
by: Chen, Hongyu, et al.
Published: (2024)
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
by: Xue, Haotian, et al.
Published: (2025)
by: Xue, Haotian, et al.
Published: (2025)
FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
by: Bo, Weihao, et al.
Published: (2025)
by: Bo, Weihao, et al.
Published: (2025)
Instance Tracking in 3D Scenes from Egocentric Videos
by: Zhao, Yunhan, et al.
Published: (2023)
by: Zhao, Yunhan, et al.
Published: (2023)
Similar Items
-
Enhancing Medical Visual Grounding via Knowledge-guided Spatial Prompts
by: Gao, Yifan, et al.
Published: (2026) -
Attention to Trajectory: Trajectory-Aware Open-Vocabulary Tracking
by: Li, Yunhao, et al.
Published: (2025) -
EgoNav: Egocentric Scene-aware Human Trajectory Prediction
by: Wang, Weizhuo, et al.
Published: (2024) -
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
by: Yoshida, Tomoya, et al.
Published: (2025) -
Egocentric Event-Based Vision for Ping Pong Ball Trajectory Prediction
by: Alberico, Ivan, et al.
Published: (2025)