LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhenghao, Zhang, Ziying, Liao, Junchao, Meng, Xiangyu, Hu, Qiang, Zhu, Siyu, Zhang, Xiaoyun, Qin, Long, Wang, Weizhi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912856776114176
author Zhang, Zhenghao
Zhang, Ziying
Liao, Junchao
Meng, Xiangyu
Hu, Qiang
Zhu, Siyu
Zhang, Xiaoyun
Qin, Long
Wang, Weizhi
author_facet Zhang, Zhenghao
Zhang, Ziying
Liao, Junchao
Meng, Xiangyu
Hu, Qiang
Zhu, Siyu
Zhang, Xiaoyun
Qin, Long
Wang, Weizhi
contents Recent multimodal models for instruction-based face editing enable semantic manipulation but still struggle with precise attribute control and identity preservation. Structural facial representations such as landmarks are effective for intermediate supervision, yet most existing methods treat them as rigid geometric constraints, which can degrade identity when conditional landmarks deviate significantly from the source (e.g., large expression or pose changes, inaccurate landmark estimates). To address these limitations, we propose LaTo, a landmark-tokenized diffusion transformer for fine-grained, identity-preserving face editing. Our key innovations include: (1) a landmark tokenizer that directly quantizes raw landmark coordinates into discrete facial tokens, obviating the need for dense pixel-wise correspondence; (2) a location-mapped positional encoding and a landmark-aware classifier-free guidance that jointly facilitate flexible yet decoupled interactions among instruction, geometry, and appearance, enabling strong identity preservation; and (3) a landmark predictor that leverages vision-language models to infer target landmarks from instructions and source images, whose structured chain-of-thought improves estimation accuracy and interactive control. To mitigate data scarcity, we curate HFL-150K, to our knowledge the largest benchmark for this task, containing over 150K real face pairs with fine-grained instructions. Extensive experiments show that LaTo outperforms state-of-the-art methods by 7.8% in identity preservation and 4.6% in semantic consistency. Code and dataset will be made publicly available upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25731
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
Zhang, Zhenghao
Zhang, Ziying
Liao, Junchao
Meng, Xiangyu
Hu, Qiang
Zhu, Siyu
Zhang, Xiaoyun
Qin, Long
Wang, Weizhi
Computer Vision and Pattern Recognition
Recent multimodal models for instruction-based face editing enable semantic manipulation but still struggle with precise attribute control and identity preservation. Structural facial representations such as landmarks are effective for intermediate supervision, yet most existing methods treat them as rigid geometric constraints, which can degrade identity when conditional landmarks deviate significantly from the source (e.g., large expression or pose changes, inaccurate landmark estimates). To address these limitations, we propose LaTo, a landmark-tokenized diffusion transformer for fine-grained, identity-preserving face editing. Our key innovations include: (1) a landmark tokenizer that directly quantizes raw landmark coordinates into discrete facial tokens, obviating the need for dense pixel-wise correspondence; (2) a location-mapped positional encoding and a landmark-aware classifier-free guidance that jointly facilitate flexible yet decoupled interactions among instruction, geometry, and appearance, enabling strong identity preservation; and (3) a landmark predictor that leverages vision-language models to infer target landmarks from instructions and source images, whose structured chain-of-thought improves estimation accuracy and interactive control. To mitigate data scarcity, we curate HFL-150K, to our knowledge the largest benchmark for this task, containing over 150K real face pairs with fine-grained instructions. Extensive experiments show that LaTo outperforms state-of-the-art methods by 7.8% in identity preservation and 4.6% in semantic consistency. Code and dataset will be made publicly available upon acceptance.
title LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.25731