Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zory, Feng, Pinyuan, Wang, Bingyang, Zhao, Tianwei, Yu, Suyang, Gao, Qingying, Deng, Hokin, Ma, Ziqiao, Li, Yijiang, Luo, Dezhi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Egocentric Bias in Vision-Language Models
by: Wang, Maijunxian, et al.
Published: (2026)
by: Wang, Maijunxian, et al.
Published: (2026)
Vision Language Models Cannot Reason About Physical Transformation
by: Luo, Dezhi, et al.
Published: (2026)
by: Luo, Dezhi, et al.
Published: (2026)
Increasing Computation Resolves Conflicts in Vision Language Models
by: Wang, Bingyang, et al.
Published: (2025)
by: Wang, Bingyang, et al.
Published: (2025)
Core Knowledge Deficits in Multi-Modal Language Models
by: Li, Yijiang, et al.
Published: (2024)
by: Li, Yijiang, et al.
Published: (2024)
Rethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World Modelling
by: Luo, Dezhi, et al.
Published: (2025)
by: Luo, Dezhi, et al.
Published: (2025)
Probing Mechanical Reasoning in Large Vision Language Models
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
Vision Language Models See What You Want but not What You See
by: Gao, Qingying, et al.
Published: (2024)
by: Gao, Qingying, et al.
Published: (2024)
Vision Language Models Know Law of Conservation without Understanding More-or-Less
by: Luo, Dezhi, et al.
Published: (2024)
by: Luo, Dezhi, et al.
Published: (2024)
Probing Perceptual Constancy in Large Vision-Language Models
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
The Philosophical Foundations of Growing AI Like A Child
by: Luo, Dezhi, et al.
Published: (2025)
by: Luo, Dezhi, et al.
Published: (2025)
Accessible Nonverbal Cues to Support Conversations in VR for Blind and Low Vision People
by: Jung, Crescentia, et al.
Published: (2024)
by: Jung, Crescentia, et al.
Published: (2024)
Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
by: Deng, Hokin
Published: (2025)
by: Deng, Hokin
Published: (2025)
Reading Between the Lines: How Electronic Nonverbal Cues shape Emotion Decoding
by: Kumar, Taara, et al.
Published: (2026)
by: Kumar, Taara, et al.
Published: (2026)
On SkipGram Word Embedding Models with Negative Sampling: Unified Framework and Impact of Noise Distributions
by: Liu, Dezhi, et al.
Published: (2020)
by: Liu, Dezhi, et al.
Published: (2020)
Gaze Prediction in Virtual Reality Without Eye Tracking Using Visual and Head Motion Cues
by: Petrou, Christos, et al.
Published: (2026)
by: Petrou, Christos, et al.
Published: (2026)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
by: Mazzamuto, Michele, et al.
Published: (2024)
by: Mazzamuto, Michele, et al.
Published: (2024)
Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation
by: Ma, Ziqiao, et al.
Published: (2025)
by: Ma, Ziqiao, et al.
Published: (2025)
Initiation of Interaction Detection Framework using a Nonverbal Cue for Human-Robot Interaction
by: Yun, Guhnoo, et al.
Published: (2026)
by: Yun, Guhnoo, et al.
Published: (2026)
GazeDETR: Gaze Detection using Disentangled Head and Gaze Representations
by: de Belen, Ryan Anthony Jalova, et al.
Published: (2025)
by: de Belen, Ryan Anthony Jalova, et al.
Published: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
Better artificial intelligence does not mean better models of biology
by: Linsley, Drew, et al.
Published: (2025)
by: Linsley, Drew, et al.
Published: (2025)
Anger Speaks Louder? Exploring the Effects of AI Nonverbal Emotional Cues on Human Decision Certainty in Moral Dilemmas
by: Zhang, Chenyi, et al.
Published: (2024)
by: Zhang, Chenyi, et al.
Published: (2024)
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models
by: Ma, Ziqiao, et al.
Published: (2023)
by: Ma, Ziqiao, et al.
Published: (2023)
Automatic Replication of LLM Mistakes in Medical Conversations
by: Proniakin, Oleksii, et al.
Published: (2025)
by: Proniakin, Oleksii, et al.
Published: (2025)
HAGI++: Head-Assisted Gaze Imputation and Generation
by: Jiao, Chuhan, et al.
Published: (2025)
by: Jiao, Chuhan, et al.
Published: (2025)
Enhancing Human-Robot Interaction in Healthcare: A Study on Nonverbal Communication Cues and Trust Dynamics with NAO Robot Caregivers
by: Raju, S M Taslim Uddin
Published: (2025)
by: Raju, S M Taslim Uddin
Published: (2025)
StyGazeTalk: Learning Stylized Generation of Gaze and Head Dynamics
by: Shi, Chengwei, et al.
Published: (2025)
by: Shi, Chengwei, et al.
Published: (2025)
Signaling Human Intentions to Service Robots: Understanding the Use of Social Cues during In-Person Conversations
by: Lyu, Hanfang, et al.
Published: (2025)
by: Lyu, Hanfang, et al.
Published: (2025)
GazeHTA: End-to-end Gaze Target Detection with Head-Target Association
by: Lin, Zhi-Yi, et al.
Published: (2024)
by: Lin, Zhi-Yi, et al.
Published: (2024)
Message Passing Without Temporal Direction: Constraint Semantics and the FITO Category Mistake
by: Borrill, Paul
Published: (2026)
by: Borrill, Paul
Published: (2026)
Digitalization and Virtual Assistive Systems in Tourist Mobility: Evolution, an Experience (with Observed Mistakes), Appropriate Orientations and Recommendations
by: David, Bertrand, et al.
Published: (2024)
by: David, Bertrand, et al.
Published: (2024)
DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
by: Qi, Yukun, et al.
Published: (2026)
by: Qi, Yukun, et al.
Published: (2026)
Emotional Conversation: Empowering Talking Faces with Cohesive Expression, Gaze and Pose Generation
by: Liang, Jiadong, et al.
Published: (2024)
by: Liang, Jiadong, et al.
Published: (2024)
Social Agent: Mastering Dyadic Nonverbal Behavior Generation via Conversational LLM Agents
by: Zhang, Zeyi, et al.
Published: (2025)
by: Zhang, Zeyi, et al.
Published: (2025)
TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration
by: Li, Yanshu, et al.
Published: (2025)
by: Li, Yanshu, et al.
Published: (2025)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
by: Feng, X., et al.
Published: (2024)
by: Feng, X., et al.
Published: (2024)
Nonverbal Interaction Detection
by: Wei, Jianan, et al.
Published: (2024)
by: Wei, Jianan, et al.
Published: (2024)
DynaMIC: Dynamic Multimodal In-Context Learning Enabled Embodied Robot Counterfactual Resistance Ability
by: Yan, Tianqiang, et al.
Published: (2025)
by: Yan, Tianqiang, et al.
Published: (2025)
Data-driven Head Motion Generation through Natural Gaze-Head Coordination
by: Liu, Xiaohan, et al.
Published: (2026)
by: Liu, Xiaohan, et al.
Published: (2026)
Similar Items
-
Egocentric Bias in Vision-Language Models
by: Wang, Maijunxian, et al.
Published: (2026) -
Vision Language Models Cannot Reason About Physical Transformation
by: Luo, Dezhi, et al.
Published: (2026) -
Increasing Computation Resolves Conflicts in Vision Language Models
by: Wang, Bingyang, et al.
Published: (2025) -
Core Knowledge Deficits in Multi-Modal Language Models
by: Li, Yijiang, et al.
Published: (2024) -
Rethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World Modelling
by: Luo, Dezhi, et al.
Published: (2025)