Resolving Ambiguity in Gaze-Facilitated Visual Assistant Interaction Paradigm
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zeyu, Chen, Baiyu, Yan, Kun, Piao, Hongjing, Xue, Hao, Salim, Flora D., Shi, Yuanchun, Wang, Yuntao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM Agents
by: Li, Zechen, et al.
Published: (2025)
by: Li, Zechen, et al.
Published: (2025)
Voila-A: Aligning Vision-Language Models with User's Gaze Attention
by: Yan, Kun, et al.
Published: (2023)
by: Yan, Kun, et al.
Published: (2023)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
by: Chen, Baiyu, et al.
Published: (2025)
by: Chen, Baiyu, et al.
Published: (2025)
G-VOILA: Gaze-Facilitated Information Querying in Daily Scenarios
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
Resolving Ambiguity in Composed Image Retrieval via Calibrated Interaction
by: Tran, Amsisan, et al.
Published: (2026)
by: Tran, Amsisan, et al.
Published: (2026)
Memory-efficient Low-latency Remote Photoplethysmography through Temporal-Spatial State Space Duality
by: Wang, Kegang, et al.
Published: (2025)
by: Wang, Kegang, et al.
Published: (2025)
GazeGen: Gaze-Driven User Interaction for Visual Content Generation
by: Hsieh, He-Yen, et al.
Published: (2024)
by: Hsieh, He-Yen, et al.
Published: (2024)
AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
by: Chen, Baiyu, et al.
Published: (2026)
by: Chen, Baiyu, et al.
Published: (2026)
BevSplat: Resolving Height Ambiguity via Feature-Based Gaussian Primitives for Weakly-Supervised Cross-View Localization
by: Wang, Qiwei, et al.
Published: (2025)
by: Wang, Qiwei, et al.
Published: (2025)
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
by: Jian, Pu, et al.
Published: (2025)
by: Jian, Pu, et al.
Published: (2025)
Generate the Forest before the Trees -- A Hierarchical Diffusion model for Climate Downscaling
by: Curran, Declan J., et al.
Published: (2025)
by: Curran, Declan J., et al.
Published: (2025)
CLIP-Gaze: Towards General Gaze Estimation via Visual-Linguistic Model
by: Yin, Pengwei, et al.
Published: (2024)
by: Yin, Pengwei, et al.
Published: (2024)
CLOVA: A Closed-Loop Visual Assistant with Tool Usage and Update
by: Gao, Zhi, et al.
Published: (2023)
by: Gao, Zhi, et al.
Published: (2023)
GazeCLIP: Enhancing Gaze Estimation Through Text-Guided Multimodal Learning
by: Wang, Jun, et al.
Published: (2023)
by: Wang, Jun, et al.
Published: (2023)
ViLCo-Bench: VIdeo Language COntinual learning Benchmark
by: Tang, Tianqi, et al.
Published: (2024)
by: Tang, Tianqi, et al.
Published: (2024)
Bisecle: Binding and Separation in Continual Learning for Video Language Understanding
by: Tan, Yue, et al.
Published: (2025)
by: Tan, Yue, et al.
Published: (2025)
RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language Descriptions
by: Zeng, Ziyao, et al.
Published: (2024)
by: Zeng, Ziyao, et al.
Published: (2024)
Gaze-guided Hand-Object Interaction Synthesis: Dataset and Method
by: Tian, Jie, et al.
Published: (2024)
by: Tian, Jie, et al.
Published: (2024)
Vision-based Multi-future Trajectory Prediction: A Survey
by: Huang, Renhao, et al.
Published: (2023)
by: Huang, Renhao, et al.
Published: (2023)
Resolving Symmetry Ambiguity in Correspondence-based Methods for Instance-level Object Pose Estimation
by: Lin, Yongliang, et al.
Published: (2024)
by: Lin, Yongliang, et al.
Published: (2024)
FacePhys: State of the Heart Learning
by: Wang, Kegang, et al.
Published: (2025)
by: Wang, Kegang, et al.
Published: (2025)
Summit Vitals: Multi-Camera and Multi-Signal Biosensing at High Altitudes
by: Liu, Ke, et al.
Published: (2024)
by: Liu, Ke, et al.
Published: (2024)
HiT-JEPA: A Hierarchical Self-supervised Trajectory Embedding Framework for Similarity Computation
by: Li, Lihuan, et al.
Published: (2025)
by: Li, Lihuan, et al.
Published: (2025)
Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling
by: Wu, Yihang, et al.
Published: (2026)
by: Wu, Yihang, et al.
Published: (2026)
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
by: Yao, Louie Hong, et al.
Published: (2025)
by: Yao, Louie Hong, et al.
Published: (2025)
HRVDA: High-Resolution Visual Document Assistant
by: Liu, Chaohu, et al.
Published: (2024)
by: Liu, Chaohu, et al.
Published: (2024)
Gaze-DETR: Using Expert Gaze to Reduce False Positives in Vulvovaginal Candidiasis Screening
by: Kong, Yan, et al.
Published: (2024)
by: Kong, Yan, et al.
Published: (2024)
Acknowledging Focus Ambiguity in Visual Questions
by: Chen, Chongyan, et al.
Published: (2025)
by: Chen, Chongyan, et al.
Published: (2025)
IPAdapter-Instruct: Resolving Ambiguity in Image-based Conditioning using Instruct Prompts
by: Rowles, Ciara, et al.
Published: (2024)
by: Rowles, Ciara, et al.
Published: (2024)
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
by: Mao, Jiawei, et al.
Published: (2024)
by: Mao, Jiawei, et al.
Published: (2024)
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
by: Song, Yuehao, et al.
Published: (2024)
by: Song, Yuehao, et al.
Published: (2024)
Exploring Reliable PPG Authentication on Smartwatches in Daily Scenarios
by: Tang, Jiankai, et al.
Published: (2025)
by: Tang, Jiankai, et al.
Published: (2025)
Boosting Gaze Object Prediction via Pixel-level Supervision from Vision Foundation Model
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
M3PD Dataset: Dual-view Photoplethysmography (PPG) Using Front-and-rear Cameras of Smartphones in Lab and Clinical Settings
by: Tang, Jiankai, et al.
Published: (2025)
by: Tang, Jiankai, et al.
Published: (2025)
Neural Radiance and Gaze Fields for Visual Attention Modeling in 3D Environments
by: Chubarau, Andrei, et al.
Published: (2025)
by: Chubarau, Andrei, et al.
Published: (2025)
StyGazeTalk: Learning Stylized Generation of Gaze and Head Dynamics
by: Shi, Chengwei, et al.
Published: (2025)
by: Shi, Chengwei, et al.
Published: (2025)
PINN-Cast: Exploring the Role of Continuous-Depth NODE in Transformers and Physics Informed Loss as Soft Physical Constraints in Short-term Weather Forecasting
by: Saleem, Hira, et al.
Published: (2026)
by: Saleem, Hira, et al.
Published: (2026)
GaGA: Towards Interactive Global Geolocation Assistant
by: Dou, Zhiyang, et al.
Published: (2024)
by: Dou, Zhiyang, et al.
Published: (2024)
Minimal Interaction Separated Tuning: A New Paradigm for Visual Adaptation
by: Tang, Ningyuan, et al.
Published: (2024)
by: Tang, Ningyuan, et al.
Published: (2024)
TextGaze: Gaze-Controllable Face Generation with Natural Language
by: Wang, Hengfei, et al.
Published: (2024)
by: Wang, Hengfei, et al.
Published: (2024)
Similar Items
-
ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM Agents
by: Li, Zechen, et al.
Published: (2025) -
Voila-A: Aligning Vision-Language Models with User's Gaze Attention
by: Yan, Kun, et al.
Published: (2023) -
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
by: Chen, Baiyu, et al.
Published: (2025) -
G-VOILA: Gaze-Facilitated Information Querying in Daily Scenarios
by: Wang, Zeyu, et al.
Published: (2024) -
Resolving Ambiguity in Composed Image Retrieval via Calibrated Interaction
by: Tran, Amsisan, et al.
Published: (2026)