Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Shijing, Huang, Yaping, Cui, Chaoqun, Wong, David, Cheng, Yihua, Neophytou, Alexandros, Chang, Hyung Jin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
por: Wang, Shijing, et al.
Publicado: (2025)
por: Wang, Shijing, et al.
Publicado: (2025)
TextGaze: Gaze-Controllable Face Generation with Natural Language
por: Wang, Hengfei, et al.
Publicado: (2024)
por: Wang, Hengfei, et al.
Publicado: (2024)
Suppressing Uncertainty in Gaze Estimation
por: Wang, Shijing, et al.
Publicado: (2024)
por: Wang, Shijing, et al.
Publicado: (2024)
Multi-Modal Gaze Following in Conversational Scenarios
por: Hou, Yuqi, et al.
Publicado: (2023)
por: Hou, Yuqi, et al.
Publicado: (2023)
RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single Image
por: Wang, Hengfei, et al.
Publicado: (2025)
por: Wang, Hengfei, et al.
Publicado: (2025)
What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation
por: Cheng, Yihua, et al.
Publicado: (2024)
por: Cheng, Yihua, et al.
Publicado: (2024)
See Through the Noise: Improving Domain Generalization in Gaze Estimation
por: Peng, Yanming, et al.
Publicado: (2026)
por: Peng, Yanming, et al.
Publicado: (2026)
Cross-Dataset Gaze Estimation by Evidential Inter-intra Fusion
por: Wang, Shijing, et al.
Publicado: (2024)
por: Wang, Shijing, et al.
Publicado: (2024)
Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models
por: Wang, Hengfei, et al.
Publicado: (2026)
por: Wang, Hengfei, et al.
Publicado: (2026)
3D Prior is All You Need: Cross-Task Few-shot 2D Gaze Estimation
por: Cheng, Yihua, et al.
Publicado: (2025)
por: Cheng, Yihua, et al.
Publicado: (2025)
Differential Contrastive Training for Gaze Estimation
por: Zhang, Lin, et al.
Publicado: (2025)
por: Zhang, Lin, et al.
Publicado: (2025)
GaTector+: A Unified Head-free Framework for Gaze Object and Gaze Following Prediction
por: Jin, Yang, et al.
Publicado: (2025)
por: Jin, Yang, et al.
Publicado: (2025)
Appearance-based Gaze Estimation With Deep Learning: A Review and Benchmark
por: Cheng, Yihua, et al.
Publicado: (2021)
por: Cheng, Yihua, et al.
Publicado: (2021)
Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following
por: Miao, Qiaomu, et al.
Publicado: (2024)
por: Miao, Qiaomu, et al.
Publicado: (2024)
Enhancing 3D Gaze Estimation in the Wild using Weak Supervision with Gaze Following Labels
por: Vuillecard, Pierre, et al.
Publicado: (2025)
por: Vuillecard, Pierre, et al.
Publicado: (2025)
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
por: Song, Yuehao, et al.
Publicado: (2024)
por: Song, Yuehao, et al.
Publicado: (2024)
Boosting Gaze Object Prediction via Pixel-level Supervision from Vision Foundation Model
por: Jin, Yang, et al.
Publicado: (2024)
por: Jin, Yang, et al.
Publicado: (2024)
TinyGaze: Lightweight Gaze-Gesture Recognition on Commodity Mobile Devices
por: Lei, Yaxiong, et al.
Publicado: (2026)
por: Lei, Yaxiong, et al.
Publicado: (2026)
Exploring the Zero-Shot Capabilities of Vision-Language Models for Improving Gaze Following
por: Gupta, Anshul, et al.
Publicado: (2024)
por: Gupta, Anshul, et al.
Publicado: (2024)
A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze Prediction
por: Gupta, Anshul, et al.
Publicado: (2024)
por: Gupta, Anshul, et al.
Publicado: (2024)
GazeCLIP: Enhancing Gaze Estimation Through Text-Guided Multimodal Learning
por: Wang, Jun, et al.
Publicado: (2023)
por: Wang, Jun, et al.
Publicado: (2023)
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
por: Mathew, Athul M., et al.
Publicado: (2025)
por: Mathew, Athul M., et al.
Publicado: (2025)
GazeDETR: Gaze Detection using Disentangled Head and Gaze Representations
por: de Belen, Ryan Anthony Jalova, et al.
Publicado: (2025)
por: de Belen, Ryan Anthony Jalova, et al.
Publicado: (2025)
CLIP-Gaze: Towards General Gaze Estimation via Visual-Linguistic Model
por: Yin, Pengwei, et al.
Publicado: (2024)
por: Yin, Pengwei, et al.
Publicado: (2024)
GazeGen: Gaze-Driven User Interaction for Visual Content Generation
por: Hsieh, He-Yen, et al.
Publicado: (2024)
por: Hsieh, He-Yen, et al.
Publicado: (2024)
OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild
por: Qu, Hongyu, et al.
Publicado: (2025)
por: Qu, Hongyu, et al.
Publicado: (2025)
Gaze Label Alignment: Alleviating Domain Shift for Gaze Estimation
por: Zeng, Guanzhong, et al.
Publicado: (2024)
por: Zeng, Guanzhong, et al.
Publicado: (2024)
Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach
por: Liu, Feiyang, et al.
Publicado: (2024)
por: Liu, Feiyang, et al.
Publicado: (2024)
SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze
por: Sharma, Pavan Kumar, et al.
Publicado: (2026)
por: Sharma, Pavan Kumar, et al.
Publicado: (2026)
Lightweight Gaze Estimation Model Via Fusion Global Information
por: Cheng, Zhang, et al.
Publicado: (2024)
por: Cheng, Zhang, et al.
Publicado: (2024)
Gaze-directed Vision GNN for Mitigating Shortcut Learning in Medical Image
por: Wu, Shaoxuan, et al.
Publicado: (2024)
por: Wu, Shaoxuan, et al.
Publicado: (2024)
GazeShift: Unsupervised Gaze Estimation and Dataset for VR
por: Shapira, Gil, et al.
Publicado: (2026)
por: Shapira, Gil, et al.
Publicado: (2026)
GazeMotion: Gaze-guided Human Motion Forecasting
por: Hu, Zhiming, et al.
Publicado: (2024)
por: Hu, Zhiming, et al.
Publicado: (2024)
Roll Your Eyes: Gaze Redirection via Explicit 3D Eyeball Rotation
por: Choi, YoungChan, et al.
Publicado: (2025)
por: Choi, YoungChan, et al.
Publicado: (2025)
Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation
por: Pani, Anupam, et al.
Publicado: (2026)
por: Pani, Anupam, et al.
Publicado: (2026)
SecureGaze: Defending Gaze Estimation Against Backdoor Attacks
por: Du, Lingyu, et al.
Publicado: (2025)
por: Du, Lingyu, et al.
Publicado: (2025)
GazeMoDiff: Gaze-guided Diffusion Model for Stochastic Human Motion Prediction
por: Yan, Haodong, et al.
Publicado: (2023)
por: Yan, Haodong, et al.
Publicado: (2023)
Multi-task Gaze Estimation Via Unidirectional Convolution
por: Cheng, Zhang, et al.
Publicado: (2024)
por: Cheng, Zhang, et al.
Publicado: (2024)
Gaze-DETR: Using Expert Gaze to Reduce False Positives in Vulvovaginal Candidiasis Screening
por: Kong, Yan, et al.
Publicado: (2024)
por: Kong, Yan, et al.
Publicado: (2024)
GazeProphet: Software-Only Gaze Prediction for VR Foveated Rendering
por: Ebadulla, Farhaan, et al.
Publicado: (2025)
por: Ebadulla, Farhaan, et al.
Publicado: (2025)
Ejemplares similares
-
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
por: Wang, Shijing, et al.
Publicado: (2025) -
TextGaze: Gaze-Controllable Face Generation with Natural Language
por: Wang, Hengfei, et al.
Publicado: (2024) -
Suppressing Uncertainty in Gaze Estimation
por: Wang, Shijing, et al.
Publicado: (2024) -
Multi-Modal Gaze Following in Conversational Scenarios
por: Hou, Yuqi, et al.
Publicado: (2023) -
RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single Image
por: Wang, Hengfei, et al.
Publicado: (2025)