Don't drop your samples! Coherence-aware training benefits Conditional diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Dufour, Nicolas, Besnier, Victor, Kalogeiton, Vicky, Picard, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation
by: Dufour, Nicolas, et al.
Published: (2024)
by: Dufour, Nicolas, et al.
Published: (2024)
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
by: Dufour, Nicolas, et al.
Published: (2025)
by: Dufour, Nicolas, et al.
Published: (2025)
Analysis of Classifier-Free Guidance Weight Schedulers
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Training-Free Synthetic Data Generation with Dual IP-Adapter Guidance
by: Boudier, Luc, et al.
Published: (2025)
by: Boudier, Luc, et al.
Published: (2025)
PoM: Efficient Image and Video Generation with the Polynomial Mixer
by: Picard, David, et al.
Published: (2024)
by: Picard, David, et al.
Published: (2024)
E.T. the Exceptional Trajectories: Text-to-camera-trajectory generation with character awareness
by: Courant, Robin, et al.
Published: (2024)
by: Courant, Robin, et al.
Published: (2024)
T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
by: Mordacq, Julie, et al.
Published: (2025)
by: Mordacq, Julie, et al.
Published: (2025)
ADAPT: Multimodal Learning for Detecting Physiological Changes under Missing Modalities
by: Mordacq, Julie, et al.
Published: (2024)
by: Mordacq, Julie, et al.
Published: (2024)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
Di$\mathtt{[M]}$O: Distilling Masked Diffusion Models into One-step Generator
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
Diffusion Reinforcement Learning via Centered Reward Distillation
by: Zhu, Yuanzhi, et al.
Published: (2026)
by: Zhu, Yuanzhi, et al.
Published: (2026)
One-step Diffusion Models with Bregman Density Ratio Matching
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
Supervised Anomaly Detection for Complex Industrial Images
by: Baitieva, Aimira, et al.
Published: (2024)
by: Baitieva, Aimira, et al.
Published: (2024)
Don't trust your eyes: on the (un)reliability of feature visualizations
by: Geirhos, Robert, et al.
Published: (2023)
by: Geirhos, Robert, et al.
Published: (2023)
Collaborating Foundation Models for Domain Generalized Semantic Segmentation
by: Benigmim, Yasser, et al.
Published: (2023)
by: Benigmim, Yasser, et al.
Published: (2023)
Don't Reach for the Stars: Rethinking Topology for Resilient Federated Learning
by: Konstantin, Mirko, et al.
Published: (2025)
by: Konstantin, Mirko, et al.
Published: (2025)
The Geometry of Noise: Why Diffusion Models Don't Need Noise Conditioning
by: Sahraee-Ardakan, Mojtaba, et al.
Published: (2026)
by: Sahraee-Ardakan, Mojtaba, et al.
Published: (2026)
Surely Large Multimodal Models (Don't) Excel in Visual Species Recognition?
by: Liu, Tian, et al.
Published: (2025)
by: Liu, Tian, et al.
Published: (2025)
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
by: Choudhury, Rohan, et al.
Published: (2024)
by: Choudhury, Rohan, et al.
Published: (2024)
How far can we go with ImageNet for Text-to-Image generation?
by: Degeorge, L., et al.
Published: (2025)
by: Degeorge, L., et al.
Published: (2025)
How to train your ViT for OOD Detection
by: Mueller, Maximilian, et al.
Published: (2024)
by: Mueller, Maximilian, et al.
Published: (2024)
Show, Don't Tell: Detecting Novel Objects by Watching Human Videos
by: Akl, James, et al.
Published: (2026)
by: Akl, James, et al.
Published: (2026)
Unified Latents (UL): How to train your latents
by: Heek, Jonathan, et al.
Published: (2026)
by: Heek, Jonathan, et al.
Published: (2026)
Pulp Motion: Framing-aware multimodal camera and human motion generation
by: Courant, Robin, et al.
Published: (2025)
by: Courant, Robin, et al.
Published: (2025)
Don't Play Favorites: Minority Guidance for Diffusion Models
by: Um, Soobin, et al.
Published: (2023)
by: Um, Soobin, et al.
Published: (2023)
You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning
by: Moutakanni, Théo, et al.
Published: (2024)
by: Moutakanni, Théo, et al.
Published: (2024)
Adapt, But Don't Forget: Fine-Tuning and Contrastive Routing for Lane Detection under Distribution Shift
by: Khan, Mohammed Abdul Hafeez, et al.
Published: (2025)
by: Khan, Mohammed Abdul Hafeez, et al.
Published: (2025)
Fast constrained sampling in pre-trained diffusion models
by: Graikos, Alexandros, et al.
Published: (2024)
by: Graikos, Alexandros, et al.
Published: (2024)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
by: Cheng, Jintao, et al.
Published: (2025)
by: Cheng, Jintao, et al.
Published: (2025)
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs
by: Janjua, Muhammad Kamran, et al.
Published: (2026)
by: Janjua, Muhammad Kamran, et al.
Published: (2026)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
by: Parmar, Mihir, et al.
Published: (2022)
by: Parmar, Mihir, et al.
Published: (2022)
Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings
by: Dhimoïla, Grégoire, et al.
Published: (2026)
by: Dhimoïla, Grégoire, et al.
Published: (2026)
Test-Time Conditioning with Representation-Aligned Visual Features
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
by: Yu, Zhuoran, et al.
Published: (2023)
by: Yu, Zhuoran, et al.
Published: (2023)
Don't Get Me Wrong: How to Apply Deep Visual Interpretations to Time Series
by: Loeffler, Christoffer, et al.
Published: (2022)
by: Loeffler, Christoffer, et al.
Published: (2022)
Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
by: Delatolas, Thanos, et al.
Published: (2025)
by: Delatolas, Thanos, et al.
Published: (2025)
Augmentation-aware Self-supervised Learning with Conditioned Projector
by: Przewięźlikowski, Marcin, et al.
Published: (2023)
by: Przewięźlikowski, Marcin, et al.
Published: (2023)
Don't Collapse Your Features: Why CenterLoss Hurts OOD Detection and Multi-Scale Mahalanobis Wins
by: Ray, Rahul D
Published: (2026)
by: Ray, Rahul D
Published: (2026)
From stability of Langevin diffusion to convergence of proximal MCMC for non-log-concave sampling
by: Renaud, Marien, et al.
Published: (2025)
by: Renaud, Marien, et al.
Published: (2025)
Similar Items
-
Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation
by: Dufour, Nicolas, et al.
Published: (2024) -
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
by: Dufour, Nicolas, et al.
Published: (2025) -
Analysis of Classifier-Free Guidance Weight Schedulers
by: Wang, Xi, et al.
Published: (2024) -
Training-Free Synthetic Data Generation with Dual IP-Adapter Guidance
by: Boudier, Luc, et al.
Published: (2025) -
PoM: Efficient Image and Video Generation with the Polynomial Mixer
by: Picard, David, et al.
Published: (2024)