Why Do Vision Language Models Struggle To Recognize Human Emotions?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agarwal, Madhav, Tsaftaris, Sotirios A., Sevilla-Lara, Laura, McDonagh, Steven |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting
von: Agarwal, Madhav, et al.
Veröffentlicht: (2025)
von: Agarwal, Madhav, et al.
Veröffentlicht: (2025)
CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models
von: Xue, Yuyang, et al.
Veröffentlicht: (2025)
von: Xue, Yuyang, et al.
Veröffentlicht: (2025)
Erase to Enhance: Data-Efficient Machine Unlearning in MRI Reconstruction
von: Xue, Yuyang, et al.
Veröffentlicht: (2024)
von: Xue, Yuyang, et al.
Veröffentlicht: (2024)
Concept-based Adversarial Attack: a Probabilistic Perspective
von: Zhang, Andi, et al.
Veröffentlicht: (2025)
von: Zhang, Andi, et al.
Veröffentlicht: (2025)
A Shift in Perspective on Causality in Domain Generalization
von: Machlanski, Damian, et al.
Veröffentlicht: (2025)
von: Machlanski, Damian, et al.
Veröffentlicht: (2025)
CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest Radiographs
von: Dutt, Raman, et al.
Veröffentlicht: (2025)
von: Dutt, Raman, et al.
Veröffentlicht: (2025)
Preventing Shortcut Learning in Medical Image Analysis through Intermediate Layer Knowledge Distillation from Specialist Teachers
von: Boland, Christopher, et al.
Veröffentlicht: (2025)
von: Boland, Christopher, et al.
Veröffentlicht: (2025)
Large Vision-Language Models as Emotion Recognizers in Context Awareness
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
Rethinking Inter-LoRA Orthogonality in Adapter Merging: Insights from Orthogonal Monte Carlo Dropout
von: Zhang, Andi, et al.
Veröffentlicht: (2025)
von: Zhang, Andi, et al.
Veröffentlicht: (2025)
MemControl: Mitigating Memorization in Diffusion Models via Automated Parameter Selection
von: Dutt, Raman, et al.
Veröffentlicht: (2024)
von: Dutt, Raman, et al.
Veröffentlicht: (2024)
SWiFT: Soft-Mask Weight Fine-tuning for Bias Mitigation
von: Yan, Junyu, et al.
Veröffentlicht: (2025)
von: Yan, Junyu, et al.
Veröffentlicht: (2025)
Causal Ordering for Structure Learning from Time Series
von: Sanchez, Pedro P., et al.
Veröffentlicht: (2025)
von: Sanchez, Pedro P., et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning for Medical Image Analysis: The Missed Opportunity
von: Dutt, Raman, et al.
Veröffentlicht: (2023)
von: Dutt, Raman, et al.
Veröffentlicht: (2023)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
A Causal Framework for Mitigating Data Shifts in Healthcare
von: Butler, Kurt, et al.
Veröffentlicht: (2026)
von: Butler, Kurt, et al.
Veröffentlicht: (2026)
Beyond Pixel Histories: World Models with Persistent 3D State
von: Garcin, Samuel, et al.
Veröffentlicht: (2026)
von: Garcin, Samuel, et al.
Veröffentlicht: (2026)
Exploiting Mixture-of-Experts Redundancy Unlocks Multimodal Generative Abilities
von: Dutt, Raman, et al.
Veröffentlicht: (2025)
von: Dutt, Raman, et al.
Veröffentlicht: (2025)
Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style
von: Limpijankit, Marvin, et al.
Veröffentlicht: (2026)
von: Limpijankit, Marvin, et al.
Veröffentlicht: (2026)
einspace: Searching for Neural Architectures from Fundamental Operations
von: Ericsson, Linus, et al.
Veröffentlicht: (2024)
von: Ericsson, Linus, et al.
Veröffentlicht: (2024)
MedVision: Dataset and Benchmark for Quantitative Medical Image Analysis
von: Yao, Yongcheng, et al.
Veröffentlicht: (2025)
von: Yao, Yongcheng, et al.
Veröffentlicht: (2025)
Causally Steered Diffusion for Automated Video Counterfactual Generation
von: Spyrou, Nikos, et al.
Veröffentlicht: (2025)
von: Spyrou, Nikos, et al.
Veröffentlicht: (2025)
Causal-Adapter: Taming Text-to-Image Diffusion for Faithful Counterfactual Generation
von: Tong, Lei, et al.
Veröffentlicht: (2025)
von: Tong, Lei, et al.
Veröffentlicht: (2025)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
von: Deng, Ken, et al.
Veröffentlicht: (2026)
von: Deng, Ken, et al.
Veröffentlicht: (2026)
Boosting Few-Shot Learning with Disentangled Self-Supervised Learning and Meta-Learning for Medical Image Classification
von: Pachetti, Eva, et al.
Veröffentlicht: (2024)
von: Pachetti, Eva, et al.
Veröffentlicht: (2024)
E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
Improving Object Detection via Local-global Contrastive Learning
von: Triantafyllidou, Danai, et al.
Veröffentlicht: (2024)
von: Triantafyllidou, Danai, et al.
Veröffentlicht: (2024)
LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
von: Kong, Fei, et al.
Veröffentlicht: (2025)
von: Kong, Fei, et al.
Veröffentlicht: (2025)
Unlocking the Potential of Weakly Labeled Data: A Co-Evolutionary Learning Framework for Abnormality Detection and Report Generation
von: Sun, Jinghan, et al.
Veröffentlicht: (2024)
von: Sun, Jinghan, et al.
Veröffentlicht: (2024)
Do Vision Language Models Understand Human Engagement in Games?
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark
von: Mushkani, Rashid
Veröffentlicht: (2025)
von: Mushkani, Rashid
Veröffentlicht: (2025)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
von: Tan, Yuwen, et al.
Veröffentlicht: (2025)
von: Tan, Yuwen, et al.
Veröffentlicht: (2025)
Do Pre-trained Vision-Language Models Encode Object States?
von: Newman, Kaleb, et al.
Veröffentlicht: (2024)
von: Newman, Kaleb, et al.
Veröffentlicht: (2024)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
Using Vision Language Models to Detect Students' Academic Emotion through Facial Expressions
von: Wang, Deliang, et al.
Veröffentlicht: (2025)
von: Wang, Deliang, et al.
Veröffentlicht: (2025)
Recognize Any Regions
von: Yang, Haosen, et al.
Veröffentlicht: (2023)
von: Yang, Haosen, et al.
Veröffentlicht: (2023)
Conformal Predictions for Human Action Recognition with Vision-Language Models
von: Tim, Bary, et al.
Veröffentlicht: (2025)
von: Tim, Bary, et al.
Veröffentlicht: (2025)
CSEval: A Framework for Evaluating Clinical Semantics in Text-to-Image Generation
von: Cronshaw, Robert, et al.
Veröffentlicht: (2026)
von: Cronshaw, Robert, et al.
Veröffentlicht: (2026)
Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models
von: Mulukutla, Vamsi Krishna, et al.
Veröffentlicht: (2025)
von: Mulukutla, Vamsi Krishna, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025) -
GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting
von: Agarwal, Madhav, et al.
Veröffentlicht: (2025) -
CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models
von: Xue, Yuyang, et al.
Veröffentlicht: (2025) -
Erase to Enhance: Data-Efficient Machine Unlearning in MRI Reconstruction
von: Xue, Yuyang, et al.
Veröffentlicht: (2024) -
Concept-based Adversarial Attack: a Probabilistic Perspective
von: Zhang, Andi, et al.
Veröffentlicht: (2025)