ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Eyring, Luca, Karthik, Shyamgopal, Roth, Karsten, Dosovitskiy, Alexey, Akata, Zeynep |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
por: Eyring, Luca, et al.
Publicado: (2025)
por: Eyring, Luca, et al.
Publicado: (2025)
Vision-by-Language for Training-Free Compositional Image Retrieval
por: Karthik, Shyamgopal, et al.
Publicado: (2023)
por: Karthik, Shyamgopal, et al.
Publicado: (2023)
Scalable Ranked Preference Optimization for Text-to-Image Generation
por: Karthik, Shyamgopal, et al.
Publicado: (2024)
por: Karthik, Shyamgopal, et al.
Publicado: (2024)
Disentangled Representation Learning with the Gromov-Monge Gap
por: Uscidda, Théo, et al.
Publicado: (2024)
por: Uscidda, Théo, et al.
Publicado: (2024)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
por: Hummel, Thomas, et al.
Publicado: (2024)
por: Hummel, Thomas, et al.
Publicado: (2024)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
por: Pach, Mateusz, et al.
Publicado: (2025)
por: Pach, Mateusz, et al.
Publicado: (2025)
ETHER: Efficient Finetuning of Large-Scale Models with Hyperplane Reflections
por: Bini, Massimo, et al.
Publicado: (2024)
por: Bini, Massimo, et al.
Publicado: (2024)
Concept-Guided Interpretability via Neural Chunking
por: Wu, Shuchen, et al.
Publicado: (2025)
por: Wu, Shuchen, et al.
Publicado: (2025)
Road Obstacle Video Segmentation
por: Rai, Shyam Nandan, et al.
Publicado: (2025)
por: Rai, Shyam Nandan, et al.
Publicado: (2025)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
por: Singhi, Nishad, et al.
Publicado: (2024)
por: Singhi, Nishad, et al.
Publicado: (2024)
Post-hoc Probabilistic Vision-Language Models
por: Baumann, Anton, et al.
Publicado: (2024)
por: Baumann, Anton, et al.
Publicado: (2024)
Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models
por: Thede, Lukas, et al.
Publicado: (2024)
por: Thede, Lukas, et al.
Publicado: (2024)
Beyond the final layer: Attentive multilayer fusion for vision transformers
por: Ciernik, Laure, et al.
Publicado: (2026)
por: Ciernik, Laure, et al.
Publicado: (2026)
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
por: Harrington, Anne, et al.
Publicado: (2025)
por: Harrington, Anne, et al.
Publicado: (2025)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
por: Girrbach, Leander, et al.
Publicado: (2025)
por: Girrbach, Leander, et al.
Publicado: (2025)
Personalizing Text-to-Image Generation to Individual Taste
por: Maerten, Anne-Sofie, et al.
Publicado: (2026)
por: Maerten, Anne-Sofie, et al.
Publicado: (2026)
Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model
por: Roth, Karsten, et al.
Publicado: (2023)
por: Roth, Karsten, et al.
Publicado: (2023)
Sparse Autoencoders are Topic Models
por: Girrbach, Leander, et al.
Publicado: (2025)
por: Girrbach, Leander, et al.
Publicado: (2025)
Unbalancedness in Neural Monge Maps Improves Unpaired Domain Translation
por: Eyring, Luca, et al.
Publicado: (2023)
por: Eyring, Luca, et al.
Publicado: (2023)
Simplifying Knowledge Transfer in Pretrained Models
por: Jain, Siddharth, et al.
Publicado: (2025)
por: Jain, Siddharth, et al.
Publicado: (2025)
Context-Aware Multimodal Pretraining
por: Roth, Karsten, et al.
Publicado: (2024)
por: Roth, Karsten, et al.
Publicado: (2024)
How to Merge Your Multimodal Models Over Time?
por: Dziadzio, Sebastian, et al.
Publicado: (2024)
por: Dziadzio, Sebastian, et al.
Publicado: (2024)
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
por: Nguyen, Viet, et al.
Publicado: (2024)
por: Nguyen, Viet, et al.
Publicado: (2024)
ImageRAGTurbo: Towards One-step Text-to-Image Generation with Retrieval-Augmented Diffusion Models
por: Qiu, Peijie, et al.
Publicado: (2026)
por: Qiu, Peijie, et al.
Publicado: (2026)
Rethinking Concept Bottleneck Models: From Pitfalls to Solutions
por: Tapli, Merve, et al.
Publicado: (2026)
por: Tapli, Merve, et al.
Publicado: (2026)
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
por: Bini, Massimo, et al.
Publicado: (2025)
por: Bini, Massimo, et al.
Publicado: (2025)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
por: Wu, Boyong, et al.
Publicado: (2026)
por: Wu, Boyong, et al.
Publicado: (2026)
Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models
por: Kurzendörfer, David, et al.
Publicado: (2024)
por: Kurzendörfer, David, et al.
Publicado: (2024)
FLAIR: VLM with Fine-grained Language-informed Image Representations
por: Xiao, Rui, et al.
Publicado: (2024)
por: Xiao, Rui, et al.
Publicado: (2024)
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
por: Ba, Ying, et al.
Publicado: (2025)
por: Ba, Ying, et al.
Publicado: (2025)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
por: Kim, Jae Myung, et al.
Publicado: (2025)
por: Kim, Jae Myung, et al.
Publicado: (2025)
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
por: Udandarao, Vishaal, et al.
Publicado: (2025)
por: Udandarao, Vishaal, et al.
Publicado: (2025)
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
por: Huang, Yiran, et al.
Publicado: (2025)
por: Huang, Yiran, et al.
Publicado: (2025)
Preliminary analysis of RGB-NIR Image Registration techniques for off-road forestry environments
por: Deoli, Pankaj, et al.
Publicado: (2026)
por: Deoli, Pankaj, et al.
Publicado: (2026)
InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization
por: Guo, Xiefan, et al.
Publicado: (2024)
por: Guo, Xiefan, et al.
Publicado: (2024)
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
por: Yang, Hongji, et al.
Publicado: (2025)
por: Yang, Hongji, et al.
Publicado: (2025)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
por: Kim, Sanghwan, et al.
Publicado: (2025)
por: Kim, Sanghwan, et al.
Publicado: (2025)
Explaining CLIP Zero-shot Predictions Through Concepts
por: Ozdemir, Onat, et al.
Publicado: (2026)
por: Ozdemir, Onat, et al.
Publicado: (2026)
A Practitioner's Guide to Continual Multimodal Pretraining
por: Roth, Karsten, et al.
Publicado: (2024)
por: Roth, Karsten, et al.
Publicado: (2024)
The Manifold Hypothesis for Gradient-Based Explanations
por: Bordt, Sebastian, et al.
Publicado: (2022)
por: Bordt, Sebastian, et al.
Publicado: (2022)
Ejemplares similares
-
Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
por: Eyring, Luca, et al.
Publicado: (2025) -
Vision-by-Language for Training-Free Compositional Image Retrieval
por: Karthik, Shyamgopal, et al.
Publicado: (2023) -
Scalable Ranked Preference Optimization for Text-to-Image Generation
por: Karthik, Shyamgopal, et al.
Publicado: (2024) -
Disentangled Representation Learning with the Gromov-Monge Gap
por: Uscidda, Théo, et al.
Publicado: (2024) -
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
por: Hummel, Thomas, et al.
Publicado: (2024)