Saved in:
Bibliographic Details
Main Authors: Zhang, Shaofeng, Chen, Xuanqi, Liao, Ning, Zhao, Haoxiang, Wang, Xiaoxing, Tan, Haoru, Wu, Sitong, Jia, Xiaosong, Fan, Qi, Yan, Junchi
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.08648
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909959169507328
author Zhang, Shaofeng
Chen, Xuanqi
Liao, Ning
Zhao, Haoxiang
Wang, Xiaoxing
Tan, Haoru
Wu, Sitong
Jia, Xiaosong
Fan, Qi
Yan, Junchi
author_facet Zhang, Shaofeng
Chen, Xuanqi
Liao, Ning
Zhao, Haoxiang
Wang, Xiaoxing
Tan, Haoru
Wu, Sitong
Jia, Xiaosong
Fan, Qi
Yan, Junchi
contents The dominance of denoising generative models (e.g., diffusion, flow-matching) in visual synthesis is tempered by their substantial training costs and inefficiencies in representation learning. While injecting discriminative representations via auxiliary alignment has proven effective, this approach still faces key limitations: the reliance on external, pre-trained encoders introduces overhead and domain shift. A dispersed-based strategy that encourages strong separation among in-batch latent representations alleviates this specific dependency. To assess the effect of the number of negative samples in generative modeling, we propose {\mname}, a plug-and-play training framework that requires no external encoders. Our method integrates a memory bank mechanism that maintains a large, dynamically updated queue of negative samples across training iterations. This decouples the number of negatives from the mini-batch size, providing abundant and high-quality negatives for a contrastive objective without a multiplicative increase in computational cost. A low-dimensional projection head is used to further minimize memory and bandwidth overhead. {\mname} offers three principal advantages: (1) it is self-contained, eliminating dependency on pretrained vision foundation models and their associated forward-pass overhead; (2) it introduces no additional parameters or computational cost during inference; and (3) it enables substantially faster convergence, achieving superior generative quality more efficiently. On ImageNet-256, {\mname} achieves a state-of-the-art FID of \textbf{2.40} within 400k steps, significantly outperforming comparable methods.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08648
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
Zhang, Shaofeng
Chen, Xuanqi
Liao, Ning
Zhao, Haoxiang
Wang, Xiaoxing
Tan, Haoru
Wu, Sitong
Jia, Xiaosong
Fan, Qi
Yan, Junchi
Computer Vision and Pattern Recognition
The dominance of denoising generative models (e.g., diffusion, flow-matching) in visual synthesis is tempered by their substantial training costs and inefficiencies in representation learning. While injecting discriminative representations via auxiliary alignment has proven effective, this approach still faces key limitations: the reliance on external, pre-trained encoders introduces overhead and domain shift. A dispersed-based strategy that encourages strong separation among in-batch latent representations alleviates this specific dependency. To assess the effect of the number of negative samples in generative modeling, we propose {\mname}, a plug-and-play training framework that requires no external encoders. Our method integrates a memory bank mechanism that maintains a large, dynamically updated queue of negative samples across training iterations. This decouples the number of negatives from the mini-batch size, providing abundant and high-quality negatives for a contrastive objective without a multiplicative increase in computational cost. A low-dimensional projection head is used to further minimize memory and bandwidth overhead. {\mname} offers three principal advantages: (1) it is self-contained, eliminating dependency on pretrained vision foundation models and their associated forward-pass overhead; (2) it introduces no additional parameters or computational cost during inference; and (3) it enables substantially faster convergence, achieving superior generative quality more efficiently. On ImageNet-256, {\mname} achieves a state-of-the-art FID of \textbf{2.40} within 400k steps, significantly outperforming comparable methods.
title Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.08648