Boosting Latent Diffusion with Perceptual Objectives
Fuente:
arXiv
Salvato in:
| Autori principali: | Berrada, Tariq, Astolfi, Pietro, Hall, Melissa, Havasi, Marton, Benchetrit, Yohann, Romero-Soriano, Adriana, Alahari, Karteek, Drozdzal, Michal, Verbeek, Jakob |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On Improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2024)
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2024)
Entropy Rectifying Guidance for Diffusion and Flow Models
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2025)
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2025)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
di: Berrada, Tariq, et al.
Pubblicazione: (2023)
di: Berrada, Tariq, et al.
Pubblicazione: (2023)
EvalGIM: A Library for Evaluating Generative Image Models
di: Hall, Melissa, et al.
Pubblicazione: (2024)
di: Hall, Melissa, et al.
Pubblicazione: (2024)
Flowception: Temporally Expansive Flow Matching for Video Generation
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2025)
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2025)
Consistency-diversity-realism Pareto fronts of conditional image generative models
di: Astolfi, Pietro, et al.
Pubblicazione: (2024)
di: Astolfi, Pietro, et al.
Pubblicazione: (2024)
Object-centric Binding in Contrastive Language-Image Pretraining
di: Assouel, Rim, et al.
Pubblicazione: (2025)
di: Assouel, Rim, et al.
Pubblicazione: (2025)
Increasing the Utility of Synthetic Images through Chamfer Guidance
di: Dall'Asen, Nicola, et al.
Pubblicazione: (2025)
di: Dall'Asen, Nicola, et al.
Pubblicazione: (2025)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
di: Mañas, Oscar, et al.
Pubblicazione: (2024)
di: Mañas, Oscar, et al.
Pubblicazione: (2024)
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
di: Askari-Hemmat, Reyhane, et al.
Pubblicazione: (2025)
di: Askari-Hemmat, Reyhane, et al.
Pubblicazione: (2025)
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
di: Hall, Melissa, et al.
Pubblicazione: (2023)
di: Hall, Melissa, et al.
Pubblicazione: (2023)
Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance
di: Hemmat, Reyhane Askari, et al.
Pubblicazione: (2024)
di: Hemmat, Reyhane Askari, et al.
Pubblicazione: (2024)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
di: Hall, Melissa, et al.
Pubblicazione: (2024)
di: Hall, Melissa, et al.
Pubblicazione: (2024)
The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
di: Xiaofeng, Zhang, et al.
Pubblicazione: (2025)
di: Xiaofeng, Zhang, et al.
Pubblicazione: (2025)
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
di: Assouel, Rim, et al.
Pubblicazione: (2026)
di: Assouel, Rim, et al.
Pubblicazione: (2026)
Source-free Video Domain Adaptation by Learning from Noisy Labels
di: Dasgupta, Avijit, et al.
Pubblicazione: (2023)
di: Dasgupta, Avijit, et al.
Pubblicazione: (2023)
Inference-time Physics Alignment of Video Generative Models with Latent World Models
di: Yuan, Jianhao, et al.
Pubblicazione: (2026)
di: Yuan, Jianhao, et al.
Pubblicazione: (2026)
Multi-Modal Language Models as Text-to-Image Model Evaluators
di: Chen, Jiahui, et al.
Pubblicazione: (2025)
di: Chen, Jiahui, et al.
Pubblicazione: (2025)
Online In-Context Distillation for Low-Resource Vision Language Models
di: Kang, Zhiqi, et al.
Pubblicazione: (2025)
di: Kang, Zhiqi, et al.
Pubblicazione: (2025)
Advancing Prompt-Based Methods for Replay-Independent General Continual Learning
di: Kang, Zhiqi, et al.
Pubblicazione: (2025)
di: Kang, Zhiqi, et al.
Pubblicazione: (2025)
Exploring High-Order Self-Similarity for Video Understanding
di: Kim, Manjin, et al.
Pubblicazione: (2026)
di: Kim, Manjin, et al.
Pubblicazione: (2026)
OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
di: Nguyen, John, et al.
Pubblicazione: (2025)
di: Nguyen, John, et al.
Pubblicazione: (2025)
Evaluating the Label Efficiency of Contrastive Self-Supervised Learning for Multi-Resolution Satellite Imagery
di: Bourcier, Jules, et al.
Pubblicazione: (2022)
di: Bourcier, Jules, et al.
Pubblicazione: (2022)
On the Shift Invariance of Max Pooling Feature Maps in Convolutional Neural Networks
di: Leterme, Hubert, et al.
Pubblicazione: (2022)
di: Leterme, Hubert, et al.
Pubblicazione: (2022)
From CNNs to Shift-Invariant Twin Models Based on Complex Wavelets
di: Leterme, Hubert, et al.
Pubblicazione: (2022)
di: Leterme, Hubert, et al.
Pubblicazione: (2022)
Dynadiff: Single-stage Decoding of Images from Continuously Evolving fMRI
di: Careil, Marlène, et al.
Pubblicazione: (2025)
di: Careil, Marlène, et al.
Pubblicazione: (2025)
Feedback-guided Data Synthesis for Imbalanced Classification
di: Hemmat, Reyhane Askari, et al.
Pubblicazione: (2023)
di: Hemmat, Reyhane Askari, et al.
Pubblicazione: (2023)
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
di: Vallaeys, Théophane, et al.
Pubblicazione: (2024)
di: Vallaeys, Théophane, et al.
Pubblicazione: (2024)
DP-RDM: Adapting Diffusion Models to Private Domains Without Fine-Tuning
di: Lebensold, Jonathan, et al.
Pubblicazione: (2024)
di: Lebensold, Jonathan, et al.
Pubblicazione: (2024)
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
di: Urbanek, Jack, et al.
Pubblicazione: (2023)
di: Urbanek, Jack, et al.
Pubblicazione: (2023)
Improving the Physics of Video Generation with VJEPA-2 Reward Signal
di: Yuan, Jianhao, et al.
Pubblicazione: (2025)
di: Yuan, Jianhao, et al.
Pubblicazione: (2025)
Controlling Multimodal LLMs via Reward-guided Decoding
di: Mañas, Oscar, et al.
Pubblicazione: (2025)
di: Mañas, Oscar, et al.
Pubblicazione: (2025)
Self-Supervised Pretraining on Satellite Imagery: a Case Study on Label-Efficient Vehicle Detection
di: BOURCIER, Jules, et al.
Pubblicazione: (2022)
di: BOURCIER, Jules, et al.
Pubblicazione: (2022)
Lightweight Structure-Aware Attention for Visual Understanding
di: Kwon, Heeseung, et al.
Pubblicazione: (2022)
di: Kwon, Heeseung, et al.
Pubblicazione: (2022)
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
di: Vallaeys, Théophane, et al.
Pubblicazione: (2025)
di: Vallaeys, Théophane, et al.
Pubblicazione: (2025)
Latent Guidance in Diffusion Models for Perceptual Evaluations
di: Saini, Shreshth, et al.
Pubblicazione: (2025)
di: Saini, Shreshth, et al.
Pubblicazione: (2025)
Boosting Latent Diffusion with Flow Matching
di: Schusterbauer, Johannes, et al.
Pubblicazione: (2023)
di: Schusterbauer, Johannes, et al.
Pubblicazione: (2023)
DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models
di: Teotia, Revant, et al.
Pubblicazione: (2025)
di: Teotia, Revant, et al.
Pubblicazione: (2025)
Brain decoding: toward real-time reconstruction of visual perception
di: Benchetrit, Yohann, et al.
Pubblicazione: (2023)
di: Benchetrit, Yohann, et al.
Pubblicazione: (2023)
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
di: Agnimo, Yedidia, et al.
Pubblicazione: (2026)
di: Agnimo, Yedidia, et al.
Pubblicazione: (2026)
Documenti analoghi
-
On Improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2024) -
Entropy Rectifying Guidance for Diffusion and Flow Models
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2025) -
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
di: Berrada, Tariq, et al.
Pubblicazione: (2023) -
EvalGIM: A Library for Evaluating Generative Image Models
di: Hall, Melissa, et al.
Pubblicazione: (2024) -
Flowception: Temporally Expansive Flow Matching for Video Generation
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2025)