EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915300121772032 |
|---|---|
| author | Chua, Phoebe Fang, Cathy Mengying Ohkawa, Takehiko Kushalnagar, Raja Nanayakkara, Suranga Maes, Pattie |
| author_facet | Chua, Phoebe Fang, Cathy Mengying Ohkawa, Takehiko Kushalnagar, Raja Nanayakkara, Suranga Maes, Pattie |
| contents | Unlike spoken languages where the use of prosodic features to convey emotion is well studied, indicators of emotion in sign language remain poorly understood, creating communication barriers in critical settings. Sign languages present unique challenges as facial expressions and hand movements simultaneously serve both grammatical and emotional functions. To address this gap, we introduce EmoSign, the first sign video dataset containing sentiment and emotion labels for 200 American Sign Language (ASL) videos. We also collect open-ended descriptions of emotion cues. Annotations were done by 3 Deaf ASL signers with professional interpretation experience. Alongside the annotations, we include baseline models for sentiment and emotion classification. This dataset not only addresses a critical gap in existing sign language research but also establishes a new benchmark for understanding model capabilities in multimodal emotion recognition for sign languages. The dataset is made available at https://huggingface.co/datasets/catfang/emosign. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_17090 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language Chua, Phoebe Fang, Cathy Mengying Ohkawa, Takehiko Kushalnagar, Raja Nanayakkara, Suranga Maes, Pattie Computer Vision and Pattern Recognition Unlike spoken languages where the use of prosodic features to convey emotion is well studied, indicators of emotion in sign language remain poorly understood, creating communication barriers in critical settings. Sign languages present unique challenges as facial expressions and hand movements simultaneously serve both grammatical and emotional functions. To address this gap, we introduce EmoSign, the first sign video dataset containing sentiment and emotion labels for 200 American Sign Language (ASL) videos. We also collect open-ended descriptions of emotion cues. Annotations were done by 3 Deaf ASL signers with professional interpretation experience. Alongside the annotations, we include baseline models for sentiment and emotion classification. This dataset not only addresses a critical gap in existing sign language research but also establishes a new benchmark for understanding model capabilities in multimodal emotion recognition for sign languages. The dataset is made available at https://huggingface.co/datasets/catfang/emosign. |
| title | EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2505.17090 |