Enhancing Diffusion Face Generation with Contrastive Embeddings and SegFormer Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rawat, Dhruvraj Singh, Sherpa, Enggen, Kirupanantha, Rishikesan, Hoang, Tin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908487883161600
author Rawat, Dhruvraj Singh
Sherpa, Enggen
Kirupanantha, Rishikesan
Hoang, Tin
author_facet Rawat, Dhruvraj Singh
Sherpa, Enggen
Kirupanantha, Rishikesan
Hoang, Tin
contents We present a benchmark of diffusion models for human face generation on a small-scale CelebAMask-HQ dataset, evaluating both unconditional and conditional pipelines. Our study compares UNet and DiT architectures for unconditional generation and explores LoRA-based fine-tuning of pretrained Stable Diffusion models as a separate experiment. Building on the multi-conditioning approach of Giambi and Lisanti, which uses both attribute vectors and segmentation masks, our main contribution is the integration of an InfoNCE loss for attribute embedding and the adoption of a SegFormer-based segmentation encoder. These enhancements improve the semantic alignment and controllability of attribute-guided synthesis. Our results highlight the effectiveness of contrastive embedding learning and advanced segmentation encoding for controlled face generation in limited data settings.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Diffusion Face Generation with Contrastive Embeddings and SegFormer Guidance
Rawat, Dhruvraj Singh
Sherpa, Enggen
Kirupanantha, Rishikesan
Hoang, Tin
Computer Vision and Pattern Recognition
We present a benchmark of diffusion models for human face generation on a small-scale CelebAMask-HQ dataset, evaluating both unconditional and conditional pipelines. Our study compares UNet and DiT architectures for unconditional generation and explores LoRA-based fine-tuning of pretrained Stable Diffusion models as a separate experiment. Building on the multi-conditioning approach of Giambi and Lisanti, which uses both attribute vectors and segmentation masks, our main contribution is the integration of an InfoNCE loss for attribute embedding and the adoption of a SegFormer-based segmentation encoder. These enhancements improve the semantic alignment and controllability of attribute-guided synthesis. Our results highlight the effectiveness of contrastive embedding learning and advanced segmentation encoding for controlled face generation in limited data settings.
title Enhancing Diffusion Face Generation with Contrastive Embeddings and SegFormer Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.09847