Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909333827092480 |
|---|---|
| author | Dehghani, Mahshid Shafiee, Amirahmad Shafiei, Ali Fallah, Neda Alizadeh, Farahmand Gholinejad, Mohammad Mehdi Behroozi, Hamid Habibi, Jafar Asgari, Ehsaneddin |
| author_facet | Dehghani, Mahshid Shafiee, Amirahmad Shafiei, Ali Fallah, Neda Alizadeh, Farahmand Gholinejad, Mohammad Mehdi Behroozi, Hamid Habibi, Jafar Asgari, Ehsaneddin |
| contents | Existing 3D facial emotion modeling have been constrained by limited emotion classes and insufficient datasets. This paper introduces "Emo3D", an extensive "Text-Image-Expression dataset" spanning a wide spectrum of human emotions, each paired with images and 3D blendshapes. Leveraging Large Language Models (LLMs), we generate a diverse array of textual descriptions, facilitating the capture of a broad spectrum of emotional expressions. Using this unique dataset, we conduct a comprehensive evaluation of language-based models' fine-tuning and vision-language models like Contranstive Language Image Pretraining (CLIP) for 3D facial expression synthesis. We also introduce a new evaluation metric for this task to more directly measure the conveyed emotion. Our new evaluation metric, Emo3D, demonstrates its superiority over Mean Squared Error (MSE) metrics in assessing visual-text alignment and semantic richness in 3D facial expressions associated with human emotions. "Emo3D" has great applications in animation design, virtual reality, and emotional human-computer interaction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_02049 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description Dehghani, Mahshid Shafiee, Amirahmad Shafiei, Ali Fallah, Neda Alizadeh, Farahmand Gholinejad, Mohammad Mehdi Behroozi, Hamid Habibi, Jafar Asgari, Ehsaneddin Computer Vision and Pattern Recognition Computation and Language Graphics I.2.7; I.2.10 Existing 3D facial emotion modeling have been constrained by limited emotion classes and insufficient datasets. This paper introduces "Emo3D", an extensive "Text-Image-Expression dataset" spanning a wide spectrum of human emotions, each paired with images and 3D blendshapes. Leveraging Large Language Models (LLMs), we generate a diverse array of textual descriptions, facilitating the capture of a broad spectrum of emotional expressions. Using this unique dataset, we conduct a comprehensive evaluation of language-based models' fine-tuning and vision-language models like Contranstive Language Image Pretraining (CLIP) for 3D facial expression synthesis. We also introduce a new evaluation metric for this task to more directly measure the conveyed emotion. Our new evaluation metric, Emo3D, demonstrates its superiority over Mean Squared Error (MSE) metrics in assessing visual-text alignment and semantic richness in 3D facial expressions associated with human emotions. "Emo3D" has great applications in animation design, virtual reality, and emotional human-computer interaction. |
| title | Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description |
| topic | Computer Vision and Pattern Recognition Computation and Language Graphics I.2.7; I.2.10 |
| url | https://arxiv.org/abs/2410.02049 |