Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.08648 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909959169507328 |
|---|---|
| author | Zhang, Shaofeng Chen, Xuanqi Liao, Ning Zhao, Haoxiang Wang, Xiaoxing Tan, Haoru Wu, Sitong Jia, Xiaosong Fan, Qi Yan, Junchi |
| author_facet | Zhang, Shaofeng Chen, Xuanqi Liao, Ning Zhao, Haoxiang Wang, Xiaoxing Tan, Haoru Wu, Sitong Jia, Xiaosong Fan, Qi Yan, Junchi |
| contents | The dominance of denoising generative models (e.g., diffusion, flow-matching) in visual synthesis is tempered by their substantial training costs and inefficiencies in representation learning. While injecting discriminative representations via auxiliary alignment has proven effective, this approach still faces key limitations: the reliance on external, pre-trained encoders introduces overhead and domain shift. A dispersed-based strategy that encourages strong separation among in-batch latent representations alleviates this specific dependency. To assess the effect of the number of negative samples in generative modeling, we propose {\mname}, a plug-and-play training framework that requires no external encoders. Our method integrates a memory bank mechanism that maintains a large, dynamically updated queue of negative samples across training iterations. This decouples the number of negatives from the mini-batch size, providing abundant and high-quality negatives for a contrastive objective without a multiplicative increase in computational cost. A low-dimensional projection head is used to further minimize memory and bandwidth overhead. {\mname} offers three principal advantages: (1) it is self-contained, eliminating dependency on pretrained vision foundation models and their associated forward-pass overhead; (2) it introduces no additional parameters or computational cost during inference; and (3) it enables substantially faster convergence, achieving superior generative quality more efficiently. On ImageNet-256, {\mname} achieves a state-of-the-art FID of \textbf{2.40} within 400k steps, significantly outperforming comparable methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_08648 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank Zhang, Shaofeng Chen, Xuanqi Liao, Ning Zhao, Haoxiang Wang, Xiaoxing Tan, Haoru Wu, Sitong Jia, Xiaosong Fan, Qi Yan, Junchi Computer Vision and Pattern Recognition The dominance of denoising generative models (e.g., diffusion, flow-matching) in visual synthesis is tempered by their substantial training costs and inefficiencies in representation learning. While injecting discriminative representations via auxiliary alignment has proven effective, this approach still faces key limitations: the reliance on external, pre-trained encoders introduces overhead and domain shift. A dispersed-based strategy that encourages strong separation among in-batch latent representations alleviates this specific dependency. To assess the effect of the number of negative samples in generative modeling, we propose {\mname}, a plug-and-play training framework that requires no external encoders. Our method integrates a memory bank mechanism that maintains a large, dynamically updated queue of negative samples across training iterations. This decouples the number of negatives from the mini-batch size, providing abundant and high-quality negatives for a contrastive objective without a multiplicative increase in computational cost. A low-dimensional projection head is used to further minimize memory and bandwidth overhead. {\mname} offers three principal advantages: (1) it is self-contained, eliminating dependency on pretrained vision foundation models and their associated forward-pass overhead; (2) it introduces no additional parameters or computational cost during inference; and (3) it enables substantially faster convergence, achieving superior generative quality more efficiently. On ImageNet-256, {\mname} achieves a state-of-the-art FID of \textbf{2.40} within 400k steps, significantly outperforming comparable methods. |
| title | Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.08648 |