FIELDS: Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ling, Chen, Shi, Henglin, Kjellström, Hedvig
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915641909313536
author Ling, Chen
Shi, Henglin
Kjellström, Hedvig
author_facet Ling, Chen
Shi, Henglin
Kjellström, Hedvig
contents Facial expressions convey the bulk of emotional information in human communication, yet existing 3D face reconstruction methods often miss subtle affective details due to reliance on 2D supervision and lack of 3D ground truth. We propose FIELDS (Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision) to address these limitations by extending self-supervised 2D image consistency cues with direct 3D expression parameter supervision and an auxiliary emotion recognition branch. Our encoder is guided by authentic expression parameters from spontaneous 4D facial scans, while an intensity-aware emotion loss encourages the 3D expression parameters to capture genuine emotion content without exaggeration. This dual-supervision strategy bridges the 2D/3D domain gap and mitigates expression-intensity bias, yielding high-fidelity 3D reconstructions that preserve subtle emotional cues. From a single image, FIELDS produces emotion-rich face models with highly realistic expressions, significantly improving in-the-wild facial expression recognition performance without sacrificing naturalness.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21245
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FIELDS: Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision
Ling, Chen
Shi, Henglin
Kjellström, Hedvig
Computer Vision and Pattern Recognition
Facial expressions convey the bulk of emotional information in human communication, yet existing 3D face reconstruction methods often miss subtle affective details due to reliance on 2D supervision and lack of 3D ground truth. We propose FIELDS (Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision) to address these limitations by extending self-supervised 2D image consistency cues with direct 3D expression parameter supervision and an auxiliary emotion recognition branch. Our encoder is guided by authentic expression parameters from spontaneous 4D facial scans, while an intensity-aware emotion loss encourages the 3D expression parameters to capture genuine emotion content without exaggeration. This dual-supervision strategy bridges the 2D/3D domain gap and mitigates expression-intensity bias, yielding high-fidelity 3D reconstructions that preserve subtle emotional cues. From a single image, FIELDS produces emotion-rich face models with highly realistic expressions, significantly improving in-the-wild facial expression recognition performance without sacrificing naturalness.
title FIELDS: Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.21245