MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
Fuente:
arXiv
Saved in:
| Main Authors: | Dufour, Nicolas, Degeorge, Lucas, Ghosh, Arijit, Kalogeiton, Vicky, Picard, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How far can we go with ImageNet for Text-to-Image generation?
by: Degeorge, L., et al.
Published: (2025)
by: Degeorge, L., et al.
Published: (2025)
Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation
by: Dufour, Nicolas, et al.
Published: (2024)
by: Dufour, Nicolas, et al.
Published: (2024)
Don't drop your samples! Coherence-aware training benefits Conditional diffusion
by: Dufour, Nicolas, et al.
Published: (2024)
by: Dufour, Nicolas, et al.
Published: (2024)
One-step Diffusion Models with Bregman Density Ratio Matching
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
MIFI: MultI-camera Feature Integration for Roust 3D Distracted Driver Activity Recognition
by: Kuang, Jian, et al.
Published: (2024)
by: Kuang, Jian, et al.
Published: (2024)
E.T. the Exceptional Trajectories: Text-to-camera-trajectory generation with character awareness
by: Courant, Robin, et al.
Published: (2024)
by: Courant, Robin, et al.
Published: (2024)
Analysis of Classifier-Free Guidance Weight Schedulers
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Training-Free Synthetic Data Generation with Dual IP-Adapter Guidance
by: Boudier, Luc, et al.
Published: (2025)
by: Boudier, Luc, et al.
Published: (2025)
What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
by: Le, Minh-Quan, et al.
Published: (2025)
by: Le, Minh-Quan, et al.
Published: (2025)
Diffusion Reinforcement Learning via Centered Reward Distillation
by: Zhu, Yuanzhi, et al.
Published: (2026)
by: Zhu, Yuanzhi, et al.
Published: (2026)
PoM: Efficient Image and Video Generation with the Polynomial Mixer
by: Picard, David, et al.
Published: (2024)
by: Picard, David, et al.
Published: (2024)
Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
by: Delatolas, Thanos, et al.
Published: (2025)
by: Delatolas, Thanos, et al.
Published: (2025)
SF20K Competition 2025: Summary and findings
by: Ghermi, Ridouane, et al.
Published: (2026)
by: Ghermi, Ridouane, et al.
Published: (2026)
One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
by: Rahary, Adrien Ramanana, et al.
Published: (2026)
by: Rahary, Adrien Ramanana, et al.
Published: (2026)
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
by: Picard, David, et al.
Published: (2026)
by: Picard, David, et al.
Published: (2026)
Pulp Motion: Framing-aware multimodal camera and human motion generation
by: Courant, Robin, et al.
Published: (2025)
by: Courant, Robin, et al.
Published: (2025)
T-REGS: Minimum Spanning Tree Regularization for Self-Supervised Learning
by: Mordacq, Julie, et al.
Published: (2025)
by: Mordacq, Julie, et al.
Published: (2025)
MUSE: Manipulating Unified Framework for Synthesizing Emotions in Images via Test-Time Optimization
by: Xia, Yingjie, et al.
Published: (2025)
by: Xia, Yingjie, et al.
Published: (2025)
AKiRa: Augmentation Kit on Rays for optical video generation
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Harnessing small projectors and multiple views for efficient vision pretraining
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
Long Story Short: Story-level Video Understanding from 20K Short Films
by: Ghermi, Ridouane, et al.
Published: (2024)
by: Ghermi, Ridouane, et al.
Published: (2024)
Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
Di$\mathtt{[M]}$O: Distilling Masked Diffusion Models into One-step Generator
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
ADAPT: Multimodal Learning for Detecting Physiological Changes under Missing Modalities
by: Mordacq, Julie, et al.
Published: (2024)
by: Mordacq, Julie, et al.
Published: (2024)
MultIOD: Rehearsal-free Multihead Incremental Object Detector
by: Belouadah, Eden, et al.
Published: (2023)
by: Belouadah, Eden, et al.
Published: (2023)
Bridging Text and Image for Artist Style Transfer via Contrastive Learning
by: Liu, Zhi-Song, et al.
Published: (2024)
by: Liu, Zhi-Song, et al.
Published: (2024)
Utilizing dynamic sparsity on pretrained DETR
by: Sedghi, Reza, et al.
Published: (2025)
by: Sedghi, Reza, et al.
Published: (2025)
FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild
by: Liu, Zhi-Song, et al.
Published: (2024)
by: Liu, Zhi-Song, et al.
Published: (2024)
Collaborating Foundation Models for Domain Generalized Semantic Segmentation
by: Benigmim, Yasser, et al.
Published: (2023)
by: Benigmim, Yasser, et al.
Published: (2023)
Name Your Style: An Arbitrary Artist-aware Image Style Transfer
by: Liu, Zhi-Song, et al.
Published: (2022)
by: Liu, Zhi-Song, et al.
Published: (2022)
HP-GAN: Harnessing pretrained networks for GAN improvement with FakeTwins and discriminator consistency
by: Son, Geonhui, et al.
Published: (2026)
by: Son, Geonhui, et al.
Published: (2026)
Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation
by: Benigmim, Yasser, et al.
Published: (2025)
by: Benigmim, Yasser, et al.
Published: (2025)
MultLFG: Training-free Multi-LoRA composition using Frequency-domain Guidance
by: Roy, Aniket, et al.
Published: (2025)
by: Roy, Aniket, et al.
Published: (2025)
C-LEAD: Contrastive Learning for Enhanced Adversarial Defense
by: Ghosh, Suklav, et al.
Published: (2025)
by: Ghosh, Suklav, et al.
Published: (2025)
High-resolution efficient image generation from WiFi CSI using a pretrained latent diffusion model
by: Ramesh, Eshan, et al.
Published: (2025)
by: Ramesh, Eshan, et al.
Published: (2025)
The effectiveness of MAE pre-pretraining for billion-scale pretraining
by: Singh, Mannat, et al.
Published: (2023)
by: Singh, Mannat, et al.
Published: (2023)
LEAD: Latent Realignment for Human Motion Diffusion
by: Andreou, Nefeli, et al.
Published: (2024)
by: Andreou, Nefeli, et al.
Published: (2024)
VersaT2I: Improving Text-to-Image Models with Versatile Reward
by: Guo, Jianshu, et al.
Published: (2024)
by: Guo, Jianshu, et al.
Published: (2024)
Diverse super-resolution with pretrained deep hiererarchical VAEs
by: Prost, Jean, et al.
Published: (2022)
by: Prost, Jean, et al.
Published: (2022)
Benchmarking transferability of SSL pretraining to same and different modality segmentation tasks
by: Jiang, Jue, et al.
Published: (2026)
by: Jiang, Jue, et al.
Published: (2026)
Similar Items
-
How far can we go with ImageNet for Text-to-Image generation?
by: Degeorge, L., et al.
Published: (2025) -
Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation
by: Dufour, Nicolas, et al.
Published: (2024) -
Don't drop your samples! Coherence-aware training benefits Conditional diffusion
by: Dufour, Nicolas, et al.
Published: (2024) -
One-step Diffusion Models with Bregman Density Ratio Matching
by: Zhu, Yuanzhi, et al.
Published: (2025) -
MIFI: MultI-camera Feature Integration for Roust 3D Distracted Driver Activity Recognition
by: Kuang, Jian, et al.
Published: (2024)