A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Udandarao, Vishaal, Cherti, Mehdi, Karthik, Shyamgopal, Jitsev, Jenia, Albanie, Samuel, Bethge, Matthias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Concept-Aware Batch Sampling Improves Language-Image Pretraining
by: Ghosh, Adhiraj, et al.
Published: (2025)
by: Ghosh, Adhiraj, et al.
Published: (2025)
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
by: Ghosh, Adhiraj, et al.
Published: (2024)
by: Ghosh, Adhiraj, et al.
Published: (2024)
Solving Spatial Supersensing Without Spatial Supersensing
by: Udandarao, Vishaal, et al.
Published: (2025)
by: Udandarao, Vishaal, et al.
Published: (2025)
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025)
by: Hochlehnert, Andreas, et al.
Published: (2025)
A Practitioner's Guide to Continual Multimodal Pretraining
by: Roth, Karsten, et al.
Published: (2024)
by: Roth, Karsten, et al.
Published: (2024)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models
by: Nezhurina, Marianna, et al.
Published: (2024)
by: Nezhurina, Marianna, et al.
Published: (2024)
How to Merge Your Multimodal Models Over Time?
by: Dziadzio, Sebastian, et al.
Published: (2024)
by: Dziadzio, Sebastian, et al.
Published: (2024)
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
by: Nezhurina, Marianna, et al.
Published: (2025)
by: Nezhurina, Marianna, et al.
Published: (2025)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
CiteME: Can Language Models Accurately Cite Scientific Claims?
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
Inverse Deep Learning Ray Tracing for Heliostat Surface Prediction
by: Lewen, Jan, et al.
Published: (2024)
by: Lewen, Jan, et al.
Published: (2024)
Scalable heliostat surface predictions from focal spots: Sim-to-Real transfer of inverse Deep Learning Raytracing
by: Lewen, Jan, et al.
Published: (2025)
by: Lewen, Jan, et al.
Published: (2025)
Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play
by: Cipolina-Kun, Lucia, et al.
Published: (2025)
by: Cipolina-Kun, Lucia, et al.
Published: (2025)
Vision-by-Language for Training-Free Compositional Image Retrieval
by: Karthik, Shyamgopal, et al.
Published: (2023)
by: Karthik, Shyamgopal, et al.
Published: (2023)
Reproducible scaling laws for contrastive language-image learning
by: Cherti, Mehdi, et al.
Published: (2022)
by: Cherti, Mehdi, et al.
Published: (2022)
AudioToolAgent: An Agentic Framework for Audio-Language Models
by: Wijngaard, Gijs, et al.
Published: (2025)
by: Wijngaard, Gijs, et al.
Published: (2025)
Personalizing Text-to-Image Generation to Individual Taste
by: Maerten, Anne-Sofie, et al.
Published: (2026)
by: Maerten, Anne-Sofie, et al.
Published: (2026)
Resolving Discrepancies in Compute-Optimal Scaling of Language Models
by: Porian, Tomer, et al.
Published: (2024)
by: Porian, Tomer, et al.
Published: (2024)
Active Data Curation Effectively Distills Large-Scale Multimodal Models
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
by: Pach, Mateusz, et al.
Published: (2025)
by: Pach, Mateusz, et al.
Published: (2025)
Safeti QRA- more than just compliance
by: Cavanagh, Nic
Published: (2007)
by: Cavanagh, Nic
Published: (2007)
Mixer is more than just a model
by: Ji, Qingfeng, et al.
Published: (2024)
by: Ji, Qingfeng, et al.
Published: (2024)
A lot more than just monte!
by: Sutherland, David
Published: ()
by: Sutherland, David
Published: ()
Miss Venezuela: more than just beauty?
by: Nunzia Auletta
Published: (2013)
by: Nunzia Auletta
Published: (2013)
Improving Performance, Robustness, and Fairness of Radiographic AI Models with Finely-Controllable Synthetic Data
by: Moroianu, Stefania L., et al.
Published: (2025)
by: Moroianu, Stefania L., et al.
Published: (2025)
CREPE: Controlling Diffusion with Replica Exchange
by: He, Jiajun, et al.
Published: (2025)
by: He, Jiajun, et al.
Published: (2025)
There is more to the de Sitter horizon than just the area
by: Fischler, Willy, et al.
Published: (2024)
by: Fischler, Willy, et al.
Published: (2024)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024)
by: Hummel, Thomas, et al.
Published: (2024)
Simplifying Knowledge Transfer in Pretrained Models
by: Jain, Siddharth, et al.
Published: (2025)
by: Jain, Siddharth, et al.
Published: (2025)
GRAB: A Challenging GRaph Analysis Benchmark for Large Multimodal Models
by: Roberts, Jonathan, et al.
Published: (2024)
by: Roberts, Jonathan, et al.
Published: (2024)
Paradoxes of Social Capital
by: Cherti, Myriam
Published: (2010)
by: Cherti, Myriam
Published: (2010)
Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs
by: Schuhmann, Christoph, et al.
Published: (2025)
by: Schuhmann, Christoph, et al.
Published: (2025)
Is there a need for more than three models?
by: Stefano Noventa
Published: (2012)
by: Stefano Noventa
Published: (2012)
CREPE: Coordinate-Aware End-to-End Document Parser
by: Okamoto, Yamato, et al.
Published: (2024)
by: Okamoto, Yamato, et al.
Published: (2024)
Post-hoc Probabilistic Vision-Language Models
by: Baumann, Anton, et al.
Published: (2024)
by: Baumann, Anton, et al.
Published: (2024)
Intubation aids in hyperangulated videolaryngoscopy: essential components more than just adjuncts
by: J. Cafferkey, et al.
Published: (2024)
by: J. Cafferkey, et al.
Published: (2024)
ABC needs more than letterman / Ronald Grover
by: Grover, Ronald
by: Grover, Ronald
SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation
by: Roberts, Jonathan, et al.
Published: (2024)
by: Roberts, Jonathan, et al.
Published: (2024)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
by: Skorobogat, Ronald, et al.
Published: (2026)
by: Skorobogat, Ronald, et al.
Published: (2026)
Similar Items
-
Concept-Aware Batch Sampling Improves Language-Image Pretraining
by: Ghosh, Adhiraj, et al.
Published: (2025) -
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
by: Ghosh, Adhiraj, et al.
Published: (2024) -
Solving Spatial Supersensing Without Spatial Supersensing
by: Udandarao, Vishaal, et al.
Published: (2025) -
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025) -
A Practitioner's Guide to Continual Multimodal Pretraining
by: Roth, Karsten, et al.
Published: (2024)