On Improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Ifriqi, Tariq Berrada, Astolfi, Pietro, Hall, Melissa, Askari-Hemmat, Reyhane, Benchetrit, Yohann, Havasi, Marton, Muckley, Matthew, Alahari, Karteek, Romero-Soriano, Adriana, Verbeek, Jakob, Drozdzal, Michal |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Boosting Latent Diffusion with Perceptual Objectives
por: Berrada, Tariq, et al.
Publicado: (2024)
por: Berrada, Tariq, et al.
Publicado: (2024)
Entropy Rectifying Guidance for Diffusion and Flow Models
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2025)
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2025)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
por: Berrada, Tariq, et al.
Publicado: (2023)
por: Berrada, Tariq, et al.
Publicado: (2023)
EvalGIM: A Library for Evaluating Generative Image Models
por: Hall, Melissa, et al.
Publicado: (2024)
por: Hall, Melissa, et al.
Publicado: (2024)
Flowception: Temporally Expansive Flow Matching for Video Generation
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2025)
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2025)
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
por: Askari-Hemmat, Reyhane, et al.
Publicado: (2025)
por: Askari-Hemmat, Reyhane, et al.
Publicado: (2025)
Increasing the Utility of Synthetic Images through Chamfer Guidance
por: Dall'Asen, Nicola, et al.
Publicado: (2025)
por: Dall'Asen, Nicola, et al.
Publicado: (2025)
Consistency-diversity-realism Pareto fronts of conditional image generative models
por: Astolfi, Pietro, et al.
Publicado: (2024)
por: Astolfi, Pietro, et al.
Publicado: (2024)
Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance
por: Hemmat, Reyhane Askari, et al.
Publicado: (2024)
por: Hemmat, Reyhane Askari, et al.
Publicado: (2024)
Feedback-guided Data Synthesis for Imbalanced Classification
por: Hemmat, Reyhane Askari, et al.
Publicado: (2023)
por: Hemmat, Reyhane Askari, et al.
Publicado: (2023)
Multi-Modal Language Models as Text-to-Image Model Evaluators
por: Chen, Jiahui, et al.
Publicado: (2025)
por: Chen, Jiahui, et al.
Publicado: (2025)
Inference-time Physics Alignment of Video Generative Models with Latent World Models
por: Yuan, Jianhao, et al.
Publicado: (2026)
por: Yuan, Jianhao, et al.
Publicado: (2026)
Improving the Physics of Video Generation with VJEPA-2 Reward Signal
por: Yuan, Jianhao, et al.
Publicado: (2025)
por: Yuan, Jianhao, et al.
Publicado: (2025)
Why Less is More (Sometimes): A Theory of Data Curation
por: Dohmatob, Elvis, et al.
Publicado: (2025)
por: Dohmatob, Elvis, et al.
Publicado: (2025)
Understanding and Mitigating Tokenization Bias in Language Models
por: Phan, Buu, et al.
Publicado: (2024)
por: Phan, Buu, et al.
Publicado: (2024)
Object-centric Binding in Contrastive Language-Image Pretraining
por: Assouel, Rim, et al.
Publicado: (2025)
por: Assouel, Rim, et al.
Publicado: (2025)
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
por: Hu, Yushi, et al.
Publicado: (2025)
por: Hu, Yushi, et al.
Publicado: (2025)
OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
por: Nguyen, John, et al.
Publicado: (2025)
por: Nguyen, John, et al.
Publicado: (2025)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
por: Mañas, Oscar, et al.
Publicado: (2024)
por: Mañas, Oscar, et al.
Publicado: (2024)
Unified Text-Image Generation with Weakness-Targeted Post-Training
por: Chen, Jiahui, et al.
Publicado: (2026)
por: Chen, Jiahui, et al.
Publicado: (2026)
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
por: Phan, Buu, et al.
Publicado: (2024)
por: Phan, Buu, et al.
Publicado: (2024)
Qinco2: Vector Compression and Search with Improved Implicit Neural Codebooks
por: Vallaeys, Théophane, et al.
Publicado: (2025)
por: Vallaeys, Théophane, et al.
Publicado: (2025)
Source-free Video Domain Adaptation by Learning from Noisy Labels
por: Dasgupta, Avijit, et al.
Publicado: (2023)
por: Dasgupta, Avijit, et al.
Publicado: (2023)
Towards image compression with perfect realism at ultra-low bitrates
por: Careil, Marlène, et al.
Publicado: (2023)
por: Careil, Marlène, et al.
Publicado: (2023)
QGen: On the Ability to Generalize in Quantization Aware Training
por: AskariHemmat, MohammadHossein, et al.
Publicado: (2024)
por: AskariHemmat, MohammadHossein, et al.
Publicado: (2024)
On the Shift Invariance of Max Pooling Feature Maps in Convolutional Neural Networks
por: Leterme, Hubert, et al.
Publicado: (2022)
por: Leterme, Hubert, et al.
Publicado: (2022)
Online In-Context Distillation for Low-Resource Vision Language Models
por: Kang, Zhiqi, et al.
Publicado: (2025)
por: Kang, Zhiqi, et al.
Publicado: (2025)
From CNNs to Shift-Invariant Twin Models Based on Complex Wavelets
por: Leterme, Hubert, et al.
Publicado: (2022)
por: Leterme, Hubert, et al.
Publicado: (2022)
Advancing Prompt-Based Methods for Replay-Independent General Continual Learning
por: Kang, Zhiqi, et al.
Publicado: (2025)
por: Kang, Zhiqi, et al.
Publicado: (2025)
Exploring High-Order Self-Similarity for Video Understanding
por: Kim, Manjin, et al.
Publicado: (2026)
por: Kim, Manjin, et al.
Publicado: (2026)
Evaluating the Label Efficiency of Contrastive Self-Supervised Learning for Multi-Resolution Satellite Imagery
por: Bourcier, Jules, et al.
Publicado: (2022)
por: Bourcier, Jules, et al.
Publicado: (2022)
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
por: Hall, Melissa, et al.
Publicado: (2023)
por: Hall, Melissa, et al.
Publicado: (2023)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
por: Hall, Melissa, et al.
Publicado: (2024)
por: Hall, Melissa, et al.
Publicado: (2024)
Dynadiff: Single-stage Decoding of Images from Continuously Evolving fMRI
por: Careil, Marlène, et al.
Publicado: (2025)
por: Careil, Marlène, et al.
Publicado: (2025)
Brain decoding: toward real-time reconstruction of visual perception
por: Benchetrit, Yohann, et al.
Publicado: (2023)
por: Benchetrit, Yohann, et al.
Publicado: (2023)
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
por: Agnimo, Yedidia, et al.
Publicado: (2026)
por: Agnimo, Yedidia, et al.
Publicado: (2026)
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
por: Assouel, Rim, et al.
Publicado: (2026)
por: Assouel, Rim, et al.
Publicado: (2026)
The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
por: Xiaofeng, Zhang, et al.
Publicado: (2025)
por: Xiaofeng, Zhang, et al.
Publicado: (2025)
Guarantee Regions for Local Explanations
por: Havasi, Marton, et al.
Publicado: (2024)
por: Havasi, Marton, et al.
Publicado: (2024)
Diverse Concept Proposals for Concept Bottleneck Models
por: Brown, Katrina, et al.
Publicado: (2024)
por: Brown, Katrina, et al.
Publicado: (2024)
Ejemplares similares
-
Boosting Latent Diffusion with Perceptual Objectives
por: Berrada, Tariq, et al.
Publicado: (2024) -
Entropy Rectifying Guidance for Diffusion and Flow Models
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2025) -
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
por: Berrada, Tariq, et al.
Publicado: (2023) -
EvalGIM: A Library for Evaluating Generative Image Models
por: Hall, Melissa, et al.
Publicado: (2024) -
Flowception: Temporally Expansive Flow Matching for Video Generation
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2025)