Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
Fuente:
arXiv
Guardado en:
| Autores principales: | Nezhurina, Marianna, Porian, Tomer, Pucceti, Giovanni, Kerssies, Tommie, Beaumont, Romain, Cherti, Mehdi, Jitsev, Jenia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models
por: Nezhurina, Marianna, et al.
Publicado: (2024)
por: Nezhurina, Marianna, et al.
Publicado: (2024)
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
por: Udandarao, Vishaal, et al.
Publicado: (2025)
por: Udandarao, Vishaal, et al.
Publicado: (2025)
Reproducible scaling laws for contrastive language-image learning
por: Cherti, Mehdi, et al.
Publicado: (2022)
por: Cherti, Mehdi, et al.
Publicado: (2022)
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
por: Kerssies, Tommie, et al.
Publicado: (2024)
por: Kerssies, Tommie, et al.
Publicado: (2024)
Resolving Discrepancies in Compute-Optimal Scaling of Language Models
por: Porian, Tomer, et al.
Publicado: (2024)
por: Porian, Tomer, et al.
Publicado: (2024)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
por: Kerssies, Tommie, et al.
Publicado: (2024)
por: Kerssies, Tommie, et al.
Publicado: (2024)
Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play
por: Cipolina-Kun, Lucia, et al.
Publicado: (2025)
por: Cipolina-Kun, Lucia, et al.
Publicado: (2025)
Simplifying Traffic Anomaly Detection with Video Foundation Models
por: Orlova, Svetlana, et al.
Publicado: (2025)
por: Orlova, Svetlana, et al.
Publicado: (2025)
Scalable heliostat surface predictions from focal spots: Sim-to-Real transfer of inverse Deep Learning Raytracing
por: Lewen, Jan, et al.
Publicado: (2025)
por: Lewen, Jan, et al.
Publicado: (2025)
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
por: Englert, Brunó B., et al.
Publicado: (2024)
por: Englert, Brunó B., et al.
Publicado: (2024)
Concept-Aware Batch Sampling Improves Language-Image Pretraining
por: Ghosh, Adhiraj, et al.
Publicado: (2025)
por: Ghosh, Adhiraj, et al.
Publicado: (2025)
What is the Added Value of UDA in the VFM Era?
por: Englert, Brunó B., et al.
Publicado: (2025)
por: Englert, Brunó B., et al.
Publicado: (2025)
Improving Performance, Robustness, and Fairness of Radiographic AI Models with Finely-Controllable Synthetic Data
por: Moroianu, Stefania L., et al.
Publicado: (2025)
por: Moroianu, Stefania L., et al.
Publicado: (2025)
Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
por: Nezhurina, Marianna, et al.
Publicado: (2025)
por: Nezhurina, Marianna, et al.
Publicado: (2025)
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
por: Norouzi, Narges, et al.
Publicado: (2026)
por: Norouzi, Narges, et al.
Publicado: (2026)
Your ViT is Secretly an Image Segmentation Model
por: Kerssies, Tommie, et al.
Publicado: (2025)
por: Kerssies, Tommie, et al.
Publicado: (2025)
Learning in Compact Spaces with Approximately Normalized Transformer
por: Franke, Jörg K. H., et al.
Publicado: (2025)
por: Franke, Jörg K. H., et al.
Publicado: (2025)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
por: Kerssies, Tommie, et al.
Publicado: (2026)
por: Kerssies, Tommie, et al.
Publicado: (2026)
Inverse Deep Learning Ray Tracing for Heliostat Surface Prediction
por: Lewen, Jan, et al.
Publicado: (2024)
por: Lewen, Jan, et al.
Publicado: (2024)
Adversarial Robustness of Vision in Open Foundation Models
por: Fox, Jonathon, et al.
Publicado: (2025)
por: Fox, Jonathon, et al.
Publicado: (2025)
A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
por: Chen, Zhihong, et al.
Publicado: (2024)
por: Chen, Zhihong, et al.
Publicado: (2024)
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
por: Ullman, Tomer
Publicado: (2024)
por: Ullman, Tomer
Publicado: (2024)
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
por: Truong, Thanh-Dat, et al.
Publicado: (2025)
por: Truong, Thanh-Dat, et al.
Publicado: (2025)
ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models
por: Yuan, Zhenghang, et al.
Publicado: (2024)
por: Yuan, Zhenghang, et al.
Publicado: (2024)
Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset
por: Huang, Wenhui, et al.
Publicado: (2026)
por: Huang, Wenhui, et al.
Publicado: (2026)
Do Vision-Language Foundational models show Robust Visual Perception?
por: Chandhok, Shivam, et al.
Publicado: (2024)
por: Chandhok, Shivam, et al.
Publicado: (2024)
LeafNet: A Large-Scale Dataset and Comprehensive Benchmark for Foundational Vision-Language Understanding of Plant Diseases
por: Quoc, Khang Nguyen, et al.
Publicado: (2026)
por: Quoc, Khang Nguyen, et al.
Publicado: (2026)
Industrial Language-Image Dataset (ILID): Adapting Vision Foundation Models for Industrial Settings
por: Moenck, Keno, et al.
Publicado: (2024)
por: Moenck, Keno, et al.
Publicado: (2024)
Data Scaling Laws for Radiology Foundation Models
por: Ilse, Maximilian, et al.
Publicado: (2025)
por: Ilse, Maximilian, et al.
Publicado: (2025)
Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset
por: Blankemeier, Louis, et al.
Publicado: (2024)
por: Blankemeier, Louis, et al.
Publicado: (2024)
Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models
por: Farronato, Nicola, et al.
Publicado: (2026)
por: Farronato, Nicola, et al.
Publicado: (2026)
Robust SAM: On the Adversarial Robustness of Vision Foundation Models
por: Long, Jiahuan, et al.
Publicado: (2025)
por: Long, Jiahuan, et al.
Publicado: (2025)
Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling
por: Moonemans, Sander, et al.
Publicado: (2025)
por: Moonemans, Sander, et al.
Publicado: (2025)
Vision-Language Dataset Distillation
por: Wu, Xindi, et al.
Publicado: (2023)
por: Wu, Xindi, et al.
Publicado: (2023)
ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving
por: Guo, Xianda, et al.
Publicado: (2025)
por: Guo, Xianda, et al.
Publicado: (2025)
VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation
por: Sarowar, Md Selim, et al.
Publicado: (2025)
por: Sarowar, Md Selim, et al.
Publicado: (2025)
Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models
por: Ding, Xin, et al.
Publicado: (2025)
por: Ding, Xin, et al.
Publicado: (2025)
Enhancing Representation in Medical Vision-Language Foundation Models via Multi-Scale Information Extraction Techniques
por: Huang, Weijian, et al.
Publicado: (2024)
por: Huang, Weijian, et al.
Publicado: (2024)
SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
por: Perez, Alejandra, et al.
Publicado: (2025)
por: Perez, Alejandra, et al.
Publicado: (2025)
Robustness of Vision Foundation Models to Common Perturbations
por: Liu, Hongbin, et al.
Publicado: (2026)
por: Liu, Hongbin, et al.
Publicado: (2026)
Ejemplares similares
-
Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models
por: Nezhurina, Marianna, et al.
Publicado: (2024) -
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
por: Udandarao, Vishaal, et al.
Publicado: (2025) -
Reproducible scaling laws for contrastive language-image learning
por: Cherti, Mehdi, et al.
Publicado: (2022) -
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
por: Kerssies, Tommie, et al.
Publicado: (2024) -
Resolving Discrepancies in Compute-Optimal Scaling of Language Models
por: Porian, Tomer, et al.
Publicado: (2024)