Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912733548511232 |
|---|---|
| author | Tsangko, Iosif Triantafyllopoulos, Andreas Abdelmoula, Adem Mallol-Ragolta, Adria Schuller, Bjoern W. |
| author_facet | Tsangko, Iosif Triantafyllopoulos, Andreas Abdelmoula, Adem Mallol-Ragolta, Adria Schuller, Bjoern W. |
| contents | Foundation Models (FMs) are rapidly transforming Affective Computing (AC), with Vision Language Models (VLMs) now capable of recognising emotions in zero shot settings. This paper probes a critical but underexplored question: what visual cues do these models rely on to infer affect, and are these cues psychologically grounded or superficially learnt? We benchmark varying scale VLMs on a teeth annotated subset of AffectNet dataset and find consistent performance shifts depending on the presence of visible teeth. Through structured introspection of, the best-performing model, i.e., GPT-4o, we show that facial attributes like eyebrow position drive much of its affective reasoning, revealing a high degree of internal consistency in its valence-arousal predictions. These patterns highlight the emergent nature of FMs behaviour, but also reveal risks: shortcut learning, bias, and fairness issues especially in sensitive domains like mental health and education. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_19079 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition Tsangko, Iosif Triantafyllopoulos, Andreas Abdelmoula, Adem Mallol-Ragolta, Adria Schuller, Bjoern W. Computer Vision and Pattern Recognition Artificial Intelligence Human-Computer Interaction Foundation Models (FMs) are rapidly transforming Affective Computing (AC), with Vision Language Models (VLMs) now capable of recognising emotions in zero shot settings. This paper probes a critical but underexplored question: what visual cues do these models rely on to infer affect, and are these cues psychologically grounded or superficially learnt? We benchmark varying scale VLMs on a teeth annotated subset of AffectNet dataset and find consistent performance shifts depending on the presence of visible teeth. Through structured introspection of, the best-performing model, i.e., GPT-4o, we show that facial attributes like eyebrow position drive much of its affective reasoning, revealing a high degree of internal consistency in its valence-arousal predictions. These patterns highlight the emergent nature of FMs behaviour, but also reveal risks: shortcut learning, bias, and fairness issues especially in sensitive domains like mental health and education. |
| title | Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Human-Computer Interaction |
| url | https://arxiv.org/abs/2506.19079 |