Saved in:
| Main Authors: | Assouel, Rim, Astolfi, Pietro, Bordes, Florian, Drozdzal, Michal, Romero-Soriano, Adriana |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.14113 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
by: Assouel, Rim, et al.
Published: (2026)
by: Assouel, Rim, et al.
Published: (2026)
Feedback-guided Data Synthesis for Imbalanced Classification
by: Hemmat, Reyhane Askari, et al.
Published: (2023)
by: Hemmat, Reyhane Askari, et al.
Published: (2023)
Consistency-diversity-realism Pareto fronts of conditional image generative models
by: Astolfi, Pietro, et al.
Published: (2024)
by: Astolfi, Pietro, et al.
Published: (2024)
Binding Visual Features Point by Point
by: Haputhanthri, Udith, et al.
Published: (2026)
by: Haputhanthri, Udith, et al.
Published: (2026)
Multi-Modal Language Models as Text-to-Image Model Evaluators
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
Entropy Rectifying Guidance for Diffusion and Flow Models
by: Ifriqi, Tariq Berrada, et al.
Published: (2025)
by: Ifriqi, Tariq Berrada, et al.
Published: (2025)
On Improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models
by: Ifriqi, Tariq Berrada, et al.
Published: (2024)
by: Ifriqi, Tariq Berrada, et al.
Published: (2024)
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
by: Urbanek, Jack, et al.
Published: (2023)
by: Urbanek, Jack, et al.
Published: (2023)
Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance
by: Hemmat, Reyhane Askari, et al.
Published: (2024)
by: Hemmat, Reyhane Askari, et al.
Published: (2024)
Controlling Multimodal LLMs via Reward-guided Decoding
by: Mañas, Oscar, et al.
Published: (2025)
by: Mañas, Oscar, et al.
Published: (2025)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
by: Mañas, Oscar, et al.
Published: (2024)
by: Mañas, Oscar, et al.
Published: (2024)
Boosting Latent Diffusion with Perceptual Objectives
by: Berrada, Tariq, et al.
Published: (2024)
by: Berrada, Tariq, et al.
Published: (2024)
The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
by: Xiaofeng, Zhang, et al.
Published: (2025)
by: Xiaofeng, Zhang, et al.
Published: (2025)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
by: Kim, Dahun, et al.
Published: (2023)
by: Kim, Dahun, et al.
Published: (2023)
Augmented Conditioning Is Enough For Effective Training Image Generation
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
by: Song, Chull Hwan, et al.
Published: (2024)
by: Song, Chull Hwan, et al.
Published: (2024)
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
by: Hall, Melissa, et al.
Published: (2023)
by: Hall, Melissa, et al.
Published: (2023)
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
by: Bendikas, Rokas, et al.
Published: (2025)
by: Bendikas, Rokas, et al.
Published: (2025)
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
by: Zhu, Bin, et al.
Published: (2023)
by: Zhu, Bin, et al.
Published: (2023)
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
by: Askari-Hemmat, Reyhane, et al.
Published: (2025)
by: Askari-Hemmat, Reyhane, et al.
Published: (2025)
SeLIP: Similarity Enhanced Contrastive Language Image Pretraining for Multi-modal Head MRI
by: Liu, Zhiyang, et al.
Published: (2025)
by: Liu, Zhiyang, et al.
Published: (2025)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
by: Hall, Melissa, et al.
Published: (2024)
by: Hall, Melissa, et al.
Published: (2024)
Increasing the Utility of Synthetic Images through Chamfer Guidance
by: Dall'Asen, Nicola, et al.
Published: (2025)
by: Dall'Asen, Nicola, et al.
Published: (2025)
Learning Physical Dynamics for Object-centric Visual Prediction
by: Xu, Huilin, et al.
Published: (2024)
by: Xu, Huilin, et al.
Published: (2024)
Unified Text-Image Generation with Weakness-Targeted Post-Training
by: Chen, Jiahui, et al.
Published: (2026)
by: Chen, Jiahui, et al.
Published: (2026)
A Framework for Evaluating Zero-Shot Image Generation in Concept-based Explainability
by: Astolfi, Giacomo, et al.
Published: (2026)
by: Astolfi, Giacomo, et al.
Published: (2026)
Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
by: Li, Yihao, et al.
Published: (2025)
by: Li, Yihao, et al.
Published: (2025)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
by: Liu, Yufang, et al.
Published: (2024)
by: Liu, Yufang, et al.
Published: (2024)
GPS-SSL: Guided Positive Sampling to Inject Prior Into Self-Supervised Learning
by: Feizi, Aarash, et al.
Published: (2024)
by: Feizi, Aarash, et al.
Published: (2024)
Visual symbolic mechanisms: Emergent symbol processing in vision language models
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
ENCLIP: Ensembling and Clustering-Based Contrastive Language-Image Pretraining for Fashion Multimodal Search with Limited Data and Low-Quality Images
by: Naik, Prithviraj Purushottam, et al.
Published: (2024)
by: Naik, Prithviraj Purushottam, et al.
Published: (2024)
SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment
by: Lu, Wenbo
Published: (2025)
by: Lu, Wenbo
Published: (2025)
High-fidelity Person-centric Subject-to-Image Synthesis
by: Wang, Yibin, et al.
Published: (2023)
by: Wang, Yibin, et al.
Published: (2023)
Non-Contrastive Vision-Language Learning with Predictive Embedding Alignment
by: Kuhn, Lukas, et al.
Published: (2026)
by: Kuhn, Lukas, et al.
Published: (2026)
Learning from Gene Names, Expression Values and Images: Contrastive Masked Text-Image Pretraining for Spatial Transcriptomics Representation Learning
by: Qian, Jiahe, et al.
Published: (2025)
by: Qian, Jiahe, et al.
Published: (2025)
Multimodal Contrastive Pretraining of CBCT and IOS for Enhanced Tooth Segmentation
by: Son, Moo Hyun, et al.
Published: (2025)
by: Son, Moo Hyun, et al.
Published: (2025)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
by: Lavoie, Samuel, et al.
Published: (2024)
by: Lavoie, Samuel, et al.
Published: (2024)
Text-to-CT Generation via 3D Latent Diffusion Model with Contrastive Vision-Language Pretraining
by: Molino, Daniele, et al.
Published: (2025)
by: Molino, Daniele, et al.
Published: (2025)
Successes and Limitations of Object-centric Models at Compositional Generalisation
by: Montero, Milton L., et al.
Published: (2024)
by: Montero, Milton L., et al.
Published: (2024)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
by: Stilz, Florian, et al.
Published: (2026)
by: Stilz, Florian, et al.
Published: (2026)
Similar Items
-
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
by: Assouel, Rim, et al.
Published: (2026) -
Feedback-guided Data Synthesis for Imbalanced Classification
by: Hemmat, Reyhane Askari, et al.
Published: (2023) -
Consistency-diversity-realism Pareto fronts of conditional image generative models
by: Astolfi, Pietro, et al.
Published: (2024) -
Binding Visual Features Point by Point
by: Haputhanthri, Udith, et al.
Published: (2026) -
Multi-Modal Language Models as Text-to-Image Model Evaluators
by: Chen, Jiahui, et al.
Published: (2025)