SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mishra, Samarth, Saenko, Kate, Saligrama, Venkatesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
von: Mishra, Samarth, et al.
Veröffentlicht: (2023)
von: Mishra, Samarth, et al.
Veröffentlicht: (2023)
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
von: Miller, Kevin, et al.
Veröffentlicht: (2025)
von: Miller, Kevin, et al.
Veröffentlicht: (2025)
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
von: Wang, Shengao, et al.
Veröffentlicht: (2025)
von: Wang, Shengao, et al.
Veröffentlicht: (2025)
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
von: Petsiuk, Vitali, et al.
Veröffentlicht: (2024)
von: Petsiuk, Vitali, et al.
Veröffentlicht: (2024)
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models
von: Ghatkesar, Aarti, et al.
Veröffentlicht: (2025)
von: Ghatkesar, Aarti, et al.
Veröffentlicht: (2025)
Deep Companion Learning: Enhancing Generalization Through Historical Consistency
von: Zhu, Ruizhao, et al.
Veröffentlicht: (2024)
von: Zhu, Ruizhao, et al.
Veröffentlicht: (2024)
BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
von: Wang, Shengao, et al.
Veröffentlicht: (2025)
von: Wang, Shengao, et al.
Veröffentlicht: (2025)
Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization
von: Dunlop, Connor, et al.
Veröffentlicht: (2025)
von: Dunlop, Connor, et al.
Veröffentlicht: (2025)
KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models
von: Yiu, Eunice, et al.
Veröffentlicht: (2024)
von: Yiu, Eunice, et al.
Veröffentlicht: (2024)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
von: Liu, Aoming, et al.
Veröffentlicht: (2025)
von: Liu, Aoming, et al.
Veröffentlicht: (2025)
Mull-Tokens: Modality-Agnostic Latent Thinking
von: Ray, Arijit, et al.
Veröffentlicht: (2025)
von: Ray, Arijit, et al.
Veröffentlicht: (2025)
PrefGen: Multimodal Preference Learning for Preference-Conditioned Image Generation
von: Mo, Wenyi, et al.
Veröffentlicht: (2025)
von: Mo, Wenyi, et al.
Veröffentlicht: (2025)
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
von: Ogezi, Michael, et al.
Veröffentlicht: (2025)
von: Ogezi, Michael, et al.
Veröffentlicht: (2025)
TSynD: Targeted Synthetic Data Generation for Enhanced Medical Image Classification
von: Niemeijer, Joshua, et al.
Veröffentlicht: (2024)
von: Niemeijer, Joshua, et al.
Veröffentlicht: (2024)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
von: Ray, Arijit, et al.
Veröffentlicht: (2024)
von: Ray, Arijit, et al.
Veröffentlicht: (2024)
Physics-Based Benchmarking Metrics for Multimodal Synthetic Images
von: Gupta, Kishor Datta, et al.
Veröffentlicht: (2025)
von: Gupta, Kishor Datta, et al.
Veröffentlicht: (2025)
Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
von: Hu, Yunqing, et al.
Veröffentlicht: (2025)
von: Hu, Yunqing, et al.
Veröffentlicht: (2025)
Enhancing Zero-shot Personalized Image Aesthetics Assessment with Profile-aware Multimodal LLM
von: Wang, Chun, et al.
Veröffentlicht: (2026)
von: Wang, Chun, et al.
Veröffentlicht: (2026)
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
von: Huang, Weikai, et al.
Veröffentlicht: (2025)
von: Huang, Weikai, et al.
Veröffentlicht: (2025)
SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
von: Tian, Yuhao, et al.
Veröffentlicht: (2025)
von: Tian, Yuhao, et al.
Veröffentlicht: (2025)
CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference
von: Dong, Zhitong, et al.
Veröffentlicht: (2026)
von: Dong, Zhitong, et al.
Veröffentlicht: (2026)
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
von: Joshi, Soham, et al.
Veröffentlicht: (2025)
von: Joshi, Soham, et al.
Veröffentlicht: (2025)
EgoGen: An Egocentric Synthetic Data Generator
von: Li, Gen, et al.
Veröffentlicht: (2024)
von: Li, Gen, et al.
Veröffentlicht: (2024)
Depth Any Video with Scalable Synthetic Data
von: Yang, Honghui, et al.
Veröffentlicht: (2024)
von: Yang, Honghui, et al.
Veröffentlicht: (2024)
Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
von: Giakoumoglou, Nikolaos, et al.
Veröffentlicht: (2025)
von: Giakoumoglou, Nikolaos, et al.
Veröffentlicht: (2025)
PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
von: Shukla, Shreya, et al.
Veröffentlicht: (2025)
von: Shukla, Shreya, et al.
Veröffentlicht: (2025)
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
von: Zhang, Ruohong, et al.
Veröffentlicht: (2024)
von: Zhang, Ruohong, et al.
Veröffentlicht: (2024)
A Fully Transformer Based Multimodal Framework for Explainable Cancer Image Segmentation Using Radiology Reports
von: Adahada, Enobong, et al.
Veröffentlicht: (2025)
von: Adahada, Enobong, et al.
Veröffentlicht: (2025)
Evaluating Compositional Scene Understanding in Multimodal Generative Models
von: Fu, Shuhao, et al.
Veröffentlicht: (2025)
von: Fu, Shuhao, et al.
Veröffentlicht: (2025)
PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories
von: Tarrés, Gemma Canet, et al.
Veröffentlicht: (2026)
von: Tarrés, Gemma Canet, et al.
Veröffentlicht: (2026)
OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
Understanding Trade offs When Conditioning Synthetic Data
von: Trabucco, Brandon, et al.
Veröffentlicht: (2025)
von: Trabucco, Brandon, et al.
Veröffentlicht: (2025)
Synthetic Simplicity: Unveiling Bias in Medical Data Augmentation
von: Babu, Krishan Agyakari Raja, et al.
Veröffentlicht: (2024)
von: Babu, Krishan Agyakari Raja, et al.
Veröffentlicht: (2024)
Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations
von: Xue, Zhiyu, et al.
Veröffentlicht: (2025)
von: Xue, Zhiyu, et al.
Veröffentlicht: (2025)
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
von: Qraitem, Maan, et al.
Veröffentlicht: (2023)
von: Qraitem, Maan, et al.
Veröffentlicht: (2023)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
ScoreMix: Synthetic Data Generation by Score Composition in Diffusion Models Improves Recognition
von: Rahimi, Parsa, et al.
Veröffentlicht: (2025)
von: Rahimi, Parsa, et al.
Veröffentlicht: (2025)
Investigation of Accuracy and Bias in Face Recognition Trained with Synthetic Data
von: Korshunov, Pavel, et al.
Veröffentlicht: (2025)
von: Korshunov, Pavel, et al.
Veröffentlicht: (2025)
Synthetic Human Action Video Data Generation with Pose Transfer
von: Knapp, Vaclav, et al.
Veröffentlicht: (2025)
von: Knapp, Vaclav, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
von: Mishra, Samarth, et al.
Veröffentlicht: (2023) -
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
von: Miller, Kevin, et al.
Veröffentlicht: (2025) -
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
von: Wang, Shengao, et al.
Veröffentlicht: (2025) -
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
von: Petsiuk, Vitali, et al.
Veröffentlicht: (2024) -
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models
von: Ghatkesar, Aarti, et al.
Veröffentlicht: (2025)