Efficiency without Compromise: CLIP-aided Text-to-Image GANs with Increased Diversity
Fuente:
arXiv
Saved in:
| Main Authors: | Kobayashi, Yuya, Takida, Yuhta, Shibuya, Takashi, Mitsufuji, Yuki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TraSCE: Trajectory Steering for Concept Erasure
by: Jain, Anubhav, et al.
Published: (2024)
by: Jain, Anubhav, et al.
Published: (2024)
Classifier-Free Guidance inside the Attraction Basin May Cause Memorization
by: Jain, Anubhav, et al.
Published: (2024)
by: Jain, Anubhav, et al.
Published: (2024)
Denoising Multi-Beta VAE: Representation Learning for Disentanglement and Generation
by: Uppal, Anshuk, et al.
Published: (2025)
by: Uppal, Anshuk, et al.
Published: (2025)
Forging and Removing Latent-Noise Diffusion Watermarks Using a Single Image
by: Jain, Anubhav, et al.
Published: (2025)
by: Jain, Anubhav, et al.
Published: (2025)
VCT: Training Consistency Models with Variational Noise Coupling
by: Silvestri, Gianluigi, et al.
Published: (2025)
by: Silvestri, Gianluigi, et al.
Published: (2025)
SONA: Learning Conditional, Unconditional, and Mismatching-Aware Discriminator
by: Takida, Yuhta, et al.
Published: (2025)
by: Takida, Yuhta, et al.
Published: (2025)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
by: Shibuya, Takashi, et al.
Published: (2023)
by: Shibuya, Takashi, et al.
Published: (2023)
MoLA: Motion Generation and Editing with Latent Diffusion Enhanced by Adversarial Training
by: Uchida, Kengo, et al.
Published: (2024)
by: Uchida, Kengo, et al.
Published: (2024)
$\textit{Jump Your Steps}$: Optimizing Sampling Schedule of Discrete Diffusion Models
by: Park, Yong-Hyun, et al.
Published: (2024)
by: Park, Yong-Hyun, et al.
Published: (2024)
HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes
by: Takida, Yuhta, et al.
Published: (2023)
by: Takida, Yuhta, et al.
Published: (2023)
G2D2: Gradient-Guided Discrete Diffusion for Inverse Problem Solving
by: Murata, Naoki, et al.
Published: (2024)
by: Murata, Naoki, et al.
Published: (2024)
A Unified View of Score-Based and Drifting Models
by: Lai, Chieh-Hsin, et al.
Published: (2026)
by: Lai, Chieh-Hsin, et al.
Published: (2026)
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
by: Hayakawa, Akio, et al.
Published: (2025)
by: Hayakawa, Akio, et al.
Published: (2025)
PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion Teacher
by: Kim, Dongjun, et al.
Published: (2024)
by: Kim, Dongjun, et al.
Published: (2024)
MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
by: Hayakawa, Akio, et al.
Published: (2024)
by: Hayakawa, Akio, et al.
Published: (2024)
Concept-TRAK: Understanding how diffusion models learn concepts through concept-level attribution
by: Park, Yonghyun, et al.
Published: (2025)
by: Park, Yonghyun, et al.
Published: (2025)
Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion
by: Kim, Dongjun, et al.
Published: (2023)
by: Kim, Dongjun, et al.
Published: (2023)
Noise Scheduling as Information-Guided Allocation in Diffusion Training
by: Raya, Gabriel, et al.
Published: (2026)
by: Raya, Gabriel, et al.
Published: (2026)
PAVAS: Physics-Aware Video-to-Audio Synthesis
by: Hyun-Bin, Oh, et al.
Published: (2025)
by: Hyun-Bin, Oh, et al.
Published: (2025)
MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
by: Cheng, Ho Kei, et al.
Published: (2024)
by: Cheng, Ho Kei, et al.
Published: (2024)
StereoSync: Spatially-Aware Stereo Audio Generation from Video
by: Marinoni, Christian, et al.
Published: (2025)
by: Marinoni, Christian, et al.
Published: (2025)
Blind Inverse Problem Solving Made Easy by Text-to-Image Latent Diffusion
by: Dontas, Michail, et al.
Published: (2024)
by: Dontas, Michail, et al.
Published: (2024)
Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
by: Yang, Shiqi, et al.
Published: (2024)
by: Yang, Shiqi, et al.
Published: (2024)
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
by: Takahashi, Akira, et al.
Published: (2025)
by: Takahashi, Akira, et al.
Published: (2025)
Improved Object-Centric Diffusion Learning with Registers and Contrastive Alignment
by: Nguyen, Bac, et al.
Published: (2026)
by: Nguyen, Bac, et al.
Published: (2026)
Dyadic Mamba: Long-term Dyadic Human Motion Synthesis
by: Tanke, Julian, et al.
Published: (2025)
by: Tanke, Julian, et al.
Published: (2025)
MedCLIP-SAM: Bridging Text and Image Towards Universal Medical Image Segmentation
by: Koleilat, Taha, et al.
Published: (2024)
by: Koleilat, Taha, et al.
Published: (2024)
TITAN-Guide: Taming Inference-Time AligNment for Guided Text-to-Video Diffusion Models
by: Simon, Christian, et al.
Published: (2025)
by: Simon, Christian, et al.
Published: (2025)
Contrast-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment
by: Lv, Song-Lin, et al.
Published: (2025)
by: Lv, Song-Lin, et al.
Published: (2025)
Transferring Visual Explainability of Self-Explaining Models to Prediction-Only Models without Additional Training
by: Yoshikawa, Yuya, et al.
Published: (2025)
by: Yoshikawa, Yuya, et al.
Published: (2025)
Tabular GANs for uneven distribution
by: Ashrapov, Insaf
Published: (2020)
by: Ashrapov, Insaf
Published: (2020)
Intriguing Properties of Modern GANs
by: Friedman, Roy, et al.
Published: (2024)
by: Friedman, Roy, et al.
Published: (2024)
Enhancing Fingerprint Image Synthesis with GANs, Diffusion Models, and Style Transfer Techniques
by: Tang, W., et al.
Published: (2024)
by: Tang, W., et al.
Published: (2024)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
by: Saito, Koichi, et al.
Published: (2024)
by: Saito, Koichi, et al.
Published: (2024)
Backdoor Attack on Unpaired Medical Image-Text Foundation Models: A Pilot Study on MedCLIP
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
HumanGif: Single-View Human Diffusion with Generative Prior
by: Hu, Shoukang, et al.
Published: (2025)
by: Hu, Shoukang, et al.
Published: (2025)
ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
by: Yao, Xin, et al.
Published: (2025)
by: Yao, Xin, et al.
Published: (2025)
Feature Unlearning for Pre-trained GANs and VAEs
by: Moon, Saemi, et al.
Published: (2023)
by: Moon, Saemi, et al.
Published: (2023)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
Similar Items
-
TraSCE: Trajectory Steering for Concept Erasure
by: Jain, Anubhav, et al.
Published: (2024) -
Classifier-Free Guidance inside the Attraction Basin May Cause Memorization
by: Jain, Anubhav, et al.
Published: (2024) -
Denoising Multi-Beta VAE: Representation Learning for Disentanglement and Generation
by: Uppal, Anshuk, et al.
Published: (2025) -
Forging and Removing Latent-Noise Diffusion Watermarks Using a Single Image
by: Jain, Anubhav, et al.
Published: (2025) -
VCT: Training Consistency Models with Variational Noise Coupling
by: Silvestri, Gianluigi, et al.
Published: (2025)