Multi-Modal Gaze Following in Conversational Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Yuqi, Zhang, Zhongqun, Horanyi, Nora, Moon, Jaewon, Cheng, Yihua, Chang, Hyung Jin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TextGaze: Gaze-Controllable Face Generation with Natural Language
by: Wang, Hengfei, et al.
Published: (2024)
by: Wang, Hengfei, et al.
Published: (2024)
RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single Image
by: Wang, Hengfei, et al.
Published: (2025)
by: Wang, Hengfei, et al.
Published: (2025)
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
by: Wang, Shijing, et al.
Published: (2025)
by: Wang, Shijing, et al.
Published: (2025)
3D Prior is All You Need: Cross-Task Few-shot 2D Gaze Estimation
by: Cheng, Yihua, et al.
Published: (2025)
by: Cheng, Yihua, et al.
Published: (2025)
Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following
by: Wang, Shijing, et al.
Published: (2026)
by: Wang, Shijing, et al.
Published: (2026)
NL2Contact: Natural Language Guided 3D Hand-Object Contact Modeling with Diffusion Model
by: Zhang, Zhongqun, et al.
Published: (2024)
by: Zhang, Zhongqun, et al.
Published: (2024)
What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation
by: Cheng, Yihua, et al.
Published: (2024)
by: Cheng, Yihua, et al.
Published: (2024)
NCRF: Neural Contact Radiance Fields for Free-Viewpoint Rendering of Hand-Object Interaction
by: Zhang, Zhongqun, et al.
Published: (2024)
by: Zhang, Zhongqun, et al.
Published: (2024)
Appearance-based Gaze Estimation With Deep Learning: A Review and Benchmark
by: Cheng, Yihua, et al.
Published: (2021)
by: Cheng, Yihua, et al.
Published: (2021)
GaTector+: A Unified Head-free Framework for Gaze Object and Gaze Following Prediction
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze Prediction
by: Gupta, Anshul, et al.
Published: (2024)
by: Gupta, Anshul, et al.
Published: (2024)
Few Exemplar-Based General Medical Image Segmentation via Domain-Aware Selective Adaptation
by: Xu, Chen, et al.
Published: (2024)
by: Xu, Chen, et al.
Published: (2024)
Leveraging Multi-Modal Saliency and Fusion for Gaze Target Detection
by: Mathew, Athul M., et al.
Published: (2025)
by: Mathew, Athul M., et al.
Published: (2025)
Multi-task Gaze Estimation Via Unidirectional Convolution
by: Cheng, Zhang, et al.
Published: (2024)
by: Cheng, Zhang, et al.
Published: (2024)
Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios
by: Xu, Jialei, et al.
Published: (2024)
by: Xu, Jialei, et al.
Published: (2024)
Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models
by: Wang, Hengfei, et al.
Published: (2026)
by: Wang, Hengfei, et al.
Published: (2026)
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
by: Song, Yuehao, et al.
Published: (2024)
by: Song, Yuehao, et al.
Published: (2024)
Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following
by: Miao, Qiaomu, et al.
Published: (2024)
by: Miao, Qiaomu, et al.
Published: (2024)
Enhancing 3D Gaze Estimation in the Wild using Weak Supervision with Gaze Following Labels
by: Vuillecard, Pierre, et al.
Published: (2025)
by: Vuillecard, Pierre, et al.
Published: (2025)
Causal Representation-Based Domain Generalization on Gaze Estimation
by: Kim, Younghan, et al.
Published: (2024)
by: Kim, Younghan, et al.
Published: (2024)
Roll Your Eyes: Gaze Redirection via Explicit 3D Eyeball Rotation
by: Choi, YoungChan, et al.
Published: (2025)
by: Choi, YoungChan, et al.
Published: (2025)
Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach
by: Liu, Feiyang, et al.
Published: (2024)
by: Liu, Feiyang, et al.
Published: (2024)
BoIR: Box-Supervised Instance Representation for Multi-Person Pose Estimation
by: Jeong, Uyoung, et al.
Published: (2023)
by: Jeong, Uyoung, et al.
Published: (2023)
GazeGen: Gaze-Driven User Interaction for Visual Content Generation
by: Hsieh, He-Yen, et al.
Published: (2024)
by: Hsieh, He-Yen, et al.
Published: (2024)
Lightweight Gaze Estimation Model Via Fusion Global Information
by: Cheng, Zhang, et al.
Published: (2024)
by: Cheng, Zhang, et al.
Published: (2024)
Ocular Authentication: Fusion of Gaze and Periocular Modalities
by: Lohr, Dillon, et al.
Published: (2025)
by: Lohr, Dillon, et al.
Published: (2025)
Exploring the Zero-Shot Capabilities of Vision-Language Models for Improving Gaze Following
by: Gupta, Anshul, et al.
Published: (2024)
by: Gupta, Anshul, et al.
Published: (2024)
IMTBench: A Multi-Scenario Cross-Modal Collaborative Evaluation Benchmark for In-Image Machine Translation
by: Lyu, Jiahao, et al.
Published: (2026)
by: Lyu, Jiahao, et al.
Published: (2026)
Multi-view Gaze Target Estimation
by: Miao, Qiaomu, et al.
Published: (2025)
by: Miao, Qiaomu, et al.
Published: (2025)
Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge
by: Lee, Young-Jun, et al.
Published: (2024)
by: Lee, Young-Jun, et al.
Published: (2024)
Differential Contrastive Training for Gaze Estimation
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation
by: Jeong, Uyoung, et al.
Published: (2025)
by: Jeong, Uyoung, et al.
Published: (2025)
MSM-Seg: A Modality-and-Slice Memory Framework with Category-Agnostic Prompting for Multi-Modal Brain Tumor Segmentation
by: Luo, Yuxiang, et al.
Published: (2025)
by: Luo, Yuxiang, et al.
Published: (2025)
Emotional Conversation: Empowering Talking Faces with Cohesive Expression, Gaze and Pose Generation
by: Liang, Jiadong, et al.
Published: (2024)
by: Liang, Jiadong, et al.
Published: (2024)
OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
by: Yin, Bo-Wen, et al.
Published: (2025)
by: Yin, Bo-Wen, et al.
Published: (2025)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
by: Jin, Hyundong, et al.
Published: (2025)
by: Jin, Hyundong, et al.
Published: (2025)
GazeDETR: Gaze Detection using Disentangled Head and Gaze Representations
by: de Belen, Ryan Anthony Jalova, et al.
Published: (2025)
by: de Belen, Ryan Anthony Jalova, et al.
Published: (2025)
Modeling State Shifting via Local-Global Distillation for Event-Frame Gaze Tracking
by: Li, Jiading, et al.
Published: (2024)
by: Li, Jiading, et al.
Published: (2024)
PersonaBooth: Personalized Text-to-Motion Generation
by: Kim, Boeun, et al.
Published: (2025)
by: Kim, Boeun, et al.
Published: (2025)
Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
by: Qi, Zheng, et al.
Published: (2025)
by: Qi, Zheng, et al.
Published: (2025)
Similar Items
-
TextGaze: Gaze-Controllable Face Generation with Natural Language
by: Wang, Hengfei, et al.
Published: (2024) -
RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single Image
by: Wang, Hengfei, et al.
Published: (2025) -
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
by: Wang, Shijing, et al.
Published: (2025) -
3D Prior is All You Need: Cross-Task Few-shot 2D Gaze Estimation
by: Cheng, Yihua, et al.
Published: (2025) -
Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following
by: Wang, Shijing, et al.
Published: (2026)