Beyond Vision: How Large Language Models Interpret Facial Expressions from Valence-Arousal Values
Fuente:
arXiv
Salvato in:
| Autori principali: | Mehra, Vaibhav, Laban, Guy, Gunes, Hatice |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Do We Talk to Robots Like Therapists, and Do They Respond Accordingly? Language Alignment in AI Emotional Support
di: Chiang, Sophie, et al.
Pubblicazione: (2025)
di: Chiang, Sophie, et al.
Pubblicazione: (2025)
Cross-Cultural Value Awareness in Large Vision-Language Models
di: Howard, Phillip, et al.
Pubblicazione: (2026)
di: Howard, Phillip, et al.
Pubblicazione: (2026)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
di: Tumu, Akshar, et al.
Pubblicazione: (2025)
di: Tumu, Akshar, et al.
Pubblicazione: (2025)
Team RAS in 10th ABAW Competition: Multimodal Valence and Arousal Estimation Approach
di: Ryumina, Elena, et al.
Pubblicazione: (2026)
di: Ryumina, Elena, et al.
Pubblicazione: (2026)
Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
An Examination of the Compositionality of Large Generative Vision-Language Models
di: Ma, Teli, et al.
Pubblicazione: (2023)
di: Ma, Teli, et al.
Pubblicazione: (2023)
Mitigating Multilingual Hallucination in Large Vision-Language Models
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
Benchmarking Deflection and Hallucination in Large Vision-Language Models
di: Moratelli, Nicholas, et al.
Pubblicazione: (2026)
di: Moratelli, Nicholas, et al.
Pubblicazione: (2026)
Are Large Vision Language Models Good Game Players?
di: Wang, Xinyu, et al.
Pubblicazione: (2025)
di: Wang, Xinyu, et al.
Pubblicazione: (2025)
Improving Personalisation in Valence and Arousal Prediction using Data Augmentation
di: Nwadike, Munachiso, et al.
Pubblicazione: (2024)
di: Nwadike, Munachiso, et al.
Pubblicazione: (2024)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
AdaSVD: Adaptive Singular Value Decomposition for Large Language Models
di: Li, Zhiteng, et al.
Pubblicazione: (2025)
di: Li, Zhiteng, et al.
Pubblicazione: (2025)
Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling
di: Chen, Qiyuan, et al.
Pubblicazione: (2026)
di: Chen, Qiyuan, et al.
Pubblicazione: (2026)
SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
di: Chen, Yi-Chia, et al.
Pubblicazione: (2024)
di: Chen, Yi-Chia, et al.
Pubblicazione: (2024)
Using Vision Language Models to Detect Students' Academic Emotion through Facial Expressions
di: Wang, Deliang, et al.
Pubblicazione: (2025)
di: Wang, Deliang, et al.
Pubblicazione: (2025)
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
di: Huang, Han, et al.
Pubblicazione: (2024)
di: Huang, Han, et al.
Pubblicazione: (2024)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
di: Xiao, Wenyi, et al.
Pubblicazione: (2025)
di: Xiao, Wenyi, et al.
Pubblicazione: (2025)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
di: Luo, Jiayun, et al.
Pubblicazione: (2025)
di: Luo, Jiayun, et al.
Pubblicazione: (2025)
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
di: Sapkota, Ranjan, et al.
Pubblicazione: (2025)
di: Sapkota, Ranjan, et al.
Pubblicazione: (2025)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
Correctable Landmark Discovery via Large Models for Vision-Language Navigation
di: Lin, Bingqian, et al.
Pubblicazione: (2024)
di: Lin, Bingqian, et al.
Pubblicazione: (2024)
Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models
di: Xu, Shicheng, et al.
Pubblicazione: (2024)
di: Xu, Shicheng, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
di: Manevich, Avshalom, et al.
Pubblicazione: (2024)
di: Manevich, Avshalom, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
di: Mitra, Chancharik, et al.
Pubblicazione: (2024)
di: Mitra, Chancharik, et al.
Pubblicazione: (2024)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
di: Zhou, Chenyu, et al.
Pubblicazione: (2024)
di: Zhou, Chenyu, et al.
Pubblicazione: (2024)
VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
di: Xiao, Wenyi, et al.
Pubblicazione: (2026)
di: Xiao, Wenyi, et al.
Pubblicazione: (2026)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
di: Lee, Yi-Lun, et al.
Pubblicazione: (2024)
di: Lee, Yi-Lun, et al.
Pubblicazione: (2024)
Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
di: Zhao, Yuming, et al.
Pubblicazione: (2026)
di: Zhao, Yuming, et al.
Pubblicazione: (2026)
HSEmotion Team at ABAW-10 Competition: Facial Expression Recognition, Valence-Arousal Estimation, Action Unit Detection and Fine-Grained Violence Classification
di: Savchenko, Andrey V., et al.
Pubblicazione: (2026)
di: Savchenko, Andrey V., et al.
Pubblicazione: (2026)
MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network
di: Ahire, Vrushank, et al.
Pubblicazione: (2025)
di: Ahire, Vrushank, et al.
Pubblicazione: (2025)
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
di: Kim, Yoonshik, et al.
Pubblicazione: (2025)
di: Kim, Yoonshik, et al.
Pubblicazione: (2025)
How Culturally Aware are Vision-Language Models?
di: Burda-Lassen, Olena, et al.
Pubblicazione: (2024)
di: Burda-Lassen, Olena, et al.
Pubblicazione: (2024)
Beyond Symbolic Solving: Multi Chain-of-Thought Voting for Geometric Reasoning in Large Language Models
di: Siddique, Md. Abu Bakor, et al.
Pubblicazione: (2026)
di: Siddique, Md. Abu Bakor, et al.
Pubblicazione: (2026)
Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations
di: Li, Yanshu
Pubblicazione: (2025)
di: Li, Yanshu
Pubblicazione: (2025)
Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026)
di: Xiong, Guangzhi, et al.
Pubblicazione: (2026)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
di: An, Wenbin, et al.
Pubblicazione: (2024)
di: An, Wenbin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Do We Talk to Robots Like Therapists, and Do They Respond Accordingly? Language Alignment in AI Emotional Support
di: Chiang, Sophie, et al.
Pubblicazione: (2025) -
Cross-Cultural Value Awareness in Large Vision-Language Models
di: Howard, Phillip, et al.
Pubblicazione: (2026) -
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025) -
Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
di: Tumu, Akshar, et al.
Pubblicazione: (2025) -
Team RAS in 10th ABAW Competition: Multimodal Valence and Arousal Estimation Approach
di: Ryumina, Elena, et al.
Pubblicazione: (2026)