StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Seungmi, Yun, Kwan, Noh, Junyong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909737509978112
author Lee, Seungmi
Yun, Kwan
Noh, Junyong
author_facet Lee, Seungmi
Yun, Kwan
Noh, Junyong
contents We introduce StyleMM, a novel framework that can construct a stylized 3D Morphable Model (3DMM) based on user-defined text descriptions specifying a target style. Building upon a pre-trained mesh deformation network and a texture generator for original 3DMM-based realistic human faces, our approach fine-tunes these models using stylized facial images generated via text-guided image-to-image (i2i) translation with a diffusion model, which serve as stylization targets for the rendered mesh. To prevent undesired changes in identity, facial alignment, or expressions during i2i translation, we introduce a stylization method that explicitly preserves the facial attributes of the source image. By maintaining these critical attributes during image stylization, the proposed approach ensures consistent 3D style transfer across the 3DMM parameter space through image-based training. Once trained, StyleMM enables feed-forward generation of stylized face meshes with explicit control over shape, expression, and texture parameters, producing meshes with consistent vertex connectivity and animatability. Quantitative and qualitative evaluations demonstrate that our approach outperforms state-of-the-art methods in terms of identity-level facial diversity and stylization capability. The code and videos are available at [kwanyun.github.io/stylemm_page](kwanyun.github.io/stylemm_page).
format Preprint
id arxiv_https___arxiv_org_abs_2508_11203
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
Lee, Seungmi
Yun, Kwan
Noh, Junyong
Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
51-04
I.3.8; I.4.9
We introduce StyleMM, a novel framework that can construct a stylized 3D Morphable Model (3DMM) based on user-defined text descriptions specifying a target style. Building upon a pre-trained mesh deformation network and a texture generator for original 3DMM-based realistic human faces, our approach fine-tunes these models using stylized facial images generated via text-guided image-to-image (i2i) translation with a diffusion model, which serve as stylization targets for the rendered mesh. To prevent undesired changes in identity, facial alignment, or expressions during i2i translation, we introduce a stylization method that explicitly preserves the facial attributes of the source image. By maintaining these critical attributes during image stylization, the proposed approach ensures consistent 3D style transfer across the 3DMM parameter space through image-based training. Once trained, StyleMM enables feed-forward generation of stylized face meshes with explicit control over shape, expression, and texture parameters, producing meshes with consistent vertex connectivity and animatability. Quantitative and qualitative evaluations demonstrate that our approach outperforms state-of-the-art methods in terms of identity-level facial diversity and stylization capability. The code and videos are available at [kwanyun.github.io/stylemm_page](kwanyun.github.io/stylemm_page).
title StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
topic Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
51-04
I.3.8; I.4.9
url https://arxiv.org/abs/2508.11203