SynFER: Towards Boosting Facial Expression Recognition with Synthetic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Xilin, Luo, Cheng, Xian, Xiaole, Li, Bing, Khan, Muhammad Haris, Ge, Zongyuan, Xie, Weicheng, Song, Siyang, Shen, Linlin, Ghanem, Bernard, Yue, Xiangyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908484596924416
author He, Xilin
Luo, Cheng
Xian, Xiaole
Li, Bing
Khan, Muhammad Haris
Ge, Zongyuan
Xie, Weicheng
Song, Siyang
Shen, Linlin
Ghanem, Bernard
Yue, Xiangyu
author_facet He, Xilin
Luo, Cheng
Xian, Xiaole
Li, Bing
Khan, Muhammad Haris
Ge, Zongyuan
Xie, Weicheng
Song, Siyang
Shen, Linlin
Ghanem, Bernard
Yue, Xiangyu
contents Facial expression datasets remain limited in scale due to the subjectivity of annotations and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial expression analysis models, particularly foundation models, that rely on large-scale data for optimal performance. To tackle the overarching and complex challenge, instead of introducing a new large-scale dataset, we introduce SynFER (Synthesis of Facial Expressions with Refined Control), a novel synthetic framework for synthesizing facial expression image data based on high-level textual descriptions as well as more fine-grained and precise control through facial action units. To ensure the quality and reliability of the synthetic data, we propose a semantic guidance technique to steer the generation process and a pseudo-label generator to help rectify the facial expression labels for the synthetic images. To demonstrate the generation fidelity and the effectiveness of the synthetic data from SynFER, we conduct extensive experiments on representation learning using both synthetic data and real-world data. Results validate the efficacy of our approach and the synthetic data. Notably, our approach achieves a 67.23% classification accuracy on AffectNet when training solely with synthetic data equivalent to the AffectNet training set size, which increases to 69.84% when scaling up to five times the original size. Code is available here.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09865
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SynFER: Towards Boosting Facial Expression Recognition with Synthetic Data
He, Xilin
Luo, Cheng
Xian, Xiaole
Li, Bing
Khan, Muhammad Haris
Ge, Zongyuan
Xie, Weicheng
Song, Siyang
Shen, Linlin
Ghanem, Bernard
Yue, Xiangyu
Computer Vision and Pattern Recognition
Facial expression datasets remain limited in scale due to the subjectivity of annotations and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial expression analysis models, particularly foundation models, that rely on large-scale data for optimal performance. To tackle the overarching and complex challenge, instead of introducing a new large-scale dataset, we introduce SynFER (Synthesis of Facial Expressions with Refined Control), a novel synthetic framework for synthesizing facial expression image data based on high-level textual descriptions as well as more fine-grained and precise control through facial action units. To ensure the quality and reliability of the synthetic data, we propose a semantic guidance technique to steer the generation process and a pseudo-label generator to help rectify the facial expression labels for the synthetic images. To demonstrate the generation fidelity and the effectiveness of the synthetic data from SynFER, we conduct extensive experiments on representation learning using both synthetic data and real-world data. Results validate the efficacy of our approach and the synthetic data. Notably, our approach achieves a 67.23% classification accuracy on AffectNet when training solely with synthetic data equivalent to the AffectNet training set size, which increases to 69.84% when scaling up to five times the original size. Code is available here.
title SynFER: Towards Boosting Facial Expression Recognition with Synthetic Data
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.09865