SuperFace: Preference-Aligned Facial Expression Estimation Beyond Pseudo Supervision
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913098769629184 |
|---|---|
| author | Kang, Zejian Xu, Xuanyang Yang, Wentao Zheng, Kai Fei, Yuanchen Zou, Hongyuan Shan, Hui Yang, Shuo Huang, Xiangru |
| author_facet | Kang, Zejian Xu, Xuanyang Yang, Wentao Zheng, Kai Fei, Yuanchen Zou, Hongyuan Shan, Hui Yang, Shuo Huang, Xiangru |
| contents | Accurate facial estimation is crucial for realistic digital human animation, and ARKit blendshape coefficients offer an interpretable representation by mapping facial motions to semantic animation controls. However, learning high-quality ARKit coefficient prediction remains limited by the absence of reliable ground-truth supervision. Existing methods typically rely on capture software such as Live Link Face to provide pseudo labels, which may contain noisy activations, biased coefficient magnitudes, and missing or inaccurate facial actions. Consequently, models trained with supervised learning tend to reproduce imperfect pseudo labels rather than optimize for perceptual expression fidelity. In this paper, we propose SuperFace, a preference-driven framework that moves ARKit facial expression estimation from pseudo-label imitation toward human-aligned perceptual optimization. Instead of treating software-estimated coefficients as fixed ground truth, SuperFace uses them only as an initialization and further improves coefficient prediction through human preference feedback on rendered facial expressions. By aligning the model with perceptual judgments rather than numerical pseudo labels, SuperFace enables more visually faithful and expressive facial animation. Experiments show that SuperFace improves expression fidelity over Live Link Face supervision, demonstrating the effectiveness of preference-driven optimization for semantic facial action prediction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_06179 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SuperFace: Preference-Aligned Facial Expression Estimation Beyond Pseudo Supervision Kang, Zejian Xu, Xuanyang Yang, Wentao Zheng, Kai Fei, Yuanchen Zou, Hongyuan Shan, Hui Yang, Shuo Huang, Xiangru Computer Vision and Pattern Recognition Accurate facial estimation is crucial for realistic digital human animation, and ARKit blendshape coefficients offer an interpretable representation by mapping facial motions to semantic animation controls. However, learning high-quality ARKit coefficient prediction remains limited by the absence of reliable ground-truth supervision. Existing methods typically rely on capture software such as Live Link Face to provide pseudo labels, which may contain noisy activations, biased coefficient magnitudes, and missing or inaccurate facial actions. Consequently, models trained with supervised learning tend to reproduce imperfect pseudo labels rather than optimize for perceptual expression fidelity. In this paper, we propose SuperFace, a preference-driven framework that moves ARKit facial expression estimation from pseudo-label imitation toward human-aligned perceptual optimization. Instead of treating software-estimated coefficients as fixed ground truth, SuperFace uses them only as an initialization and further improves coefficient prediction through human preference feedback on rendered facial expressions. By aligning the model with perceptual judgments rather than numerical pseudo labels, SuperFace enables more visually faithful and expressive facial animation. Experiments show that SuperFace improves expression fidelity over Live Link Face supervision, demonstrating the effectiveness of preference-driven optimization for semantic facial action prediction. |
| title | SuperFace: Preference-Aligned Facial Expression Estimation Beyond Pseudo Supervision |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.06179 |