PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908940334268416 |
|---|---|
| author | Xie, Yin Chen, Zhichao Xiao, Zeyu Zhao, Yongle An, Xiang Yang, Kaicheng Ran, Zimin Guo, Jia Feng, Ziyong Deng, Jiankang |
| author_facet | Xie, Yin Chen, Zhichao Xiao, Zeyu Zhao, Yongle An, Xiang Yang, Kaicheng Ran, Zimin Guo, Jia Feng, Ziyong Deng, Jiankang |
| contents | Facial representation pre-training is crucial for tasks like facial recognition, expression analysis, and virtual reality. However, existing methods face three key challenges: (1) failing to capture distinct facial features and fine-grained semantics, (2) ignoring the spatial structure inherent to facial anatomy, and (3) inefficiently utilizing limited labeled data. To overcome these, we introduce PaCo-FR, an unsupervised framework that combines masked image modeling with patch-pixel alignment. Our approach integrates three innovative components: (1) a structured masking strategy that preserves spatial coherence by aligning with semantically meaningful facial regions, (2) a novel patch-based codebook that enhances feature discrimination with multiple candidate tokens, and (3) spatial consistency constraints that preserve geometric relationships between facial components. PaCo-FR achieves state-of-the-art performance across several facial analysis tasks with just 2 million unlabeled images for pre-training. Our method demonstrates significant improvements, particularly in scenarios with varying poses, occlusions, and lighting conditions. We believe this work advances facial representation learning and offers a scalable, efficient solution that reduces reliance on expensive annotated datasets, driving more effective facial analysis systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_09691 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training Xie, Yin Chen, Zhichao Xiao, Zeyu Zhao, Yongle An, Xiang Yang, Kaicheng Ran, Zimin Guo, Jia Feng, Ziyong Deng, Jiankang Computer Vision and Pattern Recognition Facial representation pre-training is crucial for tasks like facial recognition, expression analysis, and virtual reality. However, existing methods face three key challenges: (1) failing to capture distinct facial features and fine-grained semantics, (2) ignoring the spatial structure inherent to facial anatomy, and (3) inefficiently utilizing limited labeled data. To overcome these, we introduce PaCo-FR, an unsupervised framework that combines masked image modeling with patch-pixel alignment. Our approach integrates three innovative components: (1) a structured masking strategy that preserves spatial coherence by aligning with semantically meaningful facial regions, (2) a novel patch-based codebook that enhances feature discrimination with multiple candidate tokens, and (3) spatial consistency constraints that preserve geometric relationships between facial components. PaCo-FR achieves state-of-the-art performance across several facial analysis tasks with just 2 million unlabeled images for pre-training. Our method demonstrates significant improvements, particularly in scenarios with varying poses, occlusions, and lighting conditions. We believe this work advances facial representation learning and offers a scalable, efficient solution that reduces reliance on expensive annotated datasets, driving more effective facial analysis systems. |
| title | PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.09691 |