Precise Model Benchmarking with Only a Few Observations
Fuente:
arXiv
Saved in:
| Main Authors: | Fogliato, Riccardo, Patil, Pratik, Akpinar, Nil-Jana, Monfort, Mathew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Framework for Efficient Model Evaluation through Stratification, Sampling, and Estimation
by: Fogliato, Riccardo, et al.
Published: (2024)
by: Fogliato, Riccardo, et al.
Published: (2024)
Confidence Intervals for Error Rates in 1:1 Matching Tasks: Critical Statistical Analysis and Recommendations
by: Fogliato, Riccardo, et al.
Published: (2023)
by: Fogliato, Riccardo, et al.
Published: (2023)
Uncertainty Guarantees on Automated Precision Weeding using Conformal Prediction
by: Melki, Paul, et al.
Published: (2025)
by: Melki, Paul, et al.
Published: (2025)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
by: Rajabi, Navid, et al.
Published: (2024)
by: Rajabi, Navid, et al.
Published: (2024)
Classification of Buried Objects from Ground Penetrating Radar Images by using Second Order Deep Learning Models
by: Jafuno, Douba, et al.
Published: (2024)
by: Jafuno, Douba, et al.
Published: (2024)
CACTUS as a Reliable Tool for Early Classification of Age-related Macular Degeneration
by: Gherardini, Luca, et al.
Published: (2025)
by: Gherardini, Luca, et al.
Published: (2025)
New allometric models for the USA create a step-change in forest carbon estimation, modeling, and mapping
by: Johnson, Lucas K., et al.
Published: (2024)
by: Johnson, Lucas K., et al.
Published: (2024)
Anticipatory Understanding of Resilient Agriculture to Climate
by: Willmes, David, et al.
Published: (2024)
by: Willmes, David, et al.
Published: (2024)
Morphological Prototyping for Unsupervised Slide Representation Learning in Computational Pathology
by: Song, Andrew H., et al.
Published: (2024)
by: Song, Andrew H., et al.
Published: (2024)
Statistical Edge Detection And UDF Learning For Shape Representation
by: Foy, Virgile, et al.
Published: (2024)
by: Foy, Virgile, et al.
Published: (2024)
OpenViewer: Openness-Aware Multi-View Learning
by: Du, Shide, et al.
Published: (2024)
by: Du, Shide, et al.
Published: (2024)
Neural Fingerprints for Adversarial Attack Detection
by: Fisher, Haim, et al.
Published: (2024)
by: Fisher, Haim, et al.
Published: (2024)
Real-Time Localization and Bimodal Point Pattern Analysis of Palms Using UAV Imagery
by: Cui, Kangning, et al.
Published: (2024)
by: Cui, Kangning, et al.
Published: (2024)
Distributional Deep Learning for Super-Resolution of 4D Flow MRI under Domain Shift
by: Wen, Xiaoyi, et al.
Published: (2026)
by: Wen, Xiaoyi, et al.
Published: (2026)
A Framework for Supervised and Unsupervised Segmentation and Classification of Materials Microstructure Images
by: Zhang, Kungang, et al.
Published: (2025)
by: Zhang, Kungang, et al.
Published: (2025)
Enhancing convolutional neural network generalizability via low-rank weight approximation
by: Gao, Chenyin, et al.
Published: (2022)
by: Gao, Chenyin, et al.
Published: (2022)
Analysis of Ethnic Disparities in Autism Spectrum Disorder among Toddlers
by: Ramaharsha, Aadithya Prabha, et al.
Published: (2026)
by: Ramaharsha, Aadithya Prabha, et al.
Published: (2026)
Scalable deep fusion of spaceborne lidar and synthetic aperture radar for global forest structural complexity mapping
by: de Conto, Tiago, et al.
Published: (2025)
by: de Conto, Tiago, et al.
Published: (2025)
Visual Spatial Learning: Single-Field Spatial Interpolation Using Convolutional Neural Networks
by: Tinoco, Daniel, et al.
Published: (2026)
by: Tinoco, Daniel, et al.
Published: (2026)
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
by: Rajabi, Navid, et al.
Published: (2023)
by: Rajabi, Navid, et al.
Published: (2023)
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
by: Rajabi, Navid, et al.
Published: (2024)
by: Rajabi, Navid, et al.
Published: (2024)
SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes
by: Yang, Yifan, et al.
Published: (2025)
by: Yang, Yifan, et al.
Published: (2025)
Automatic Scoring of Cognition Drawings: Assessing the Quality of Machine-Based Scores Against a Gold Standard
by: Bethmann, Arne, et al.
Published: (2023)
by: Bethmann, Arne, et al.
Published: (2023)
On the Residual-based Neural Network for Unmodeled Distortions in Coordinate Transformation
by: Rofatto, Vinicius Francisco, et al.
Published: (2025)
by: Rofatto, Vinicius Francisco, et al.
Published: (2025)
Diffusion models for multivariate subsurface generation and efficient probabilistic inversion
by: Miele, Roberto, et al.
Published: (2025)
by: Miele, Roberto, et al.
Published: (2025)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
by: Yang, Cheng-Fu, et al.
Published: (2024)
by: Yang, Cheng-Fu, et al.
Published: (2024)
Massimo: Public Queue Monitoring and Management using Mass-Spring Model
by: Kumar, Abhijeet, et al.
Published: (2024)
by: Kumar, Abhijeet, et al.
Published: (2024)
AlphaEarth Satellite Embeddings for Modelling Climate Sensitive Diseases Towards Global Health Resilience
by: Nazir, Usman, et al.
Published: (2026)
by: Nazir, Usman, et al.
Published: (2026)
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
by: Fayyazsanavi, Pooya, et al.
Published: (2024)
by: Fayyazsanavi, Pooya, et al.
Published: (2024)
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
by: Xue, Yihao, et al.
Published: (2023)
by: Xue, Yihao, et al.
Published: (2023)
Benchmarking Vision-Language Models for French PDF-to-Markdown Conversion
by: Rigal, Bruno, et al.
Published: (2026)
by: Rigal, Bruno, et al.
Published: (2026)
AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning
by: Tang, Yuwei, et al.
Published: (2024)
by: Tang, Yuwei, et al.
Published: (2024)
Benchmarking Ultrasound Foundation Models for Fetal Plane Classification
by: Barrientos, Leya, et al.
Published: (2026)
by: Barrientos, Leya, et al.
Published: (2026)
Resolving Interference (RI): Disentangling Models for Improved Model Merging
by: Ramesh, Pratik, et al.
Published: (2026)
by: Ramesh, Pratik, et al.
Published: (2026)
Understanding Learning with Sliced-Wasserstein Requires Rethinking Informative Slices
by: Tran, Huy, et al.
Published: (2024)
by: Tran, Huy, et al.
Published: (2024)
IM-Unpack: Training and Inference with Arbitrarily Low Precision Integers
by: Zeng, Zhanpeng, et al.
Published: (2024)
by: Zeng, Zhanpeng, et al.
Published: (2024)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
Visual Planning: Let's Think Only with Images
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
by: Gafni, Tomer, et al.
Published: (2025)
by: Gafni, Tomer, et al.
Published: (2025)
Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
by: Robicheaux, Peter, et al.
Published: (2025)
by: Robicheaux, Peter, et al.
Published: (2025)
Similar Items
-
A Framework for Efficient Model Evaluation through Stratification, Sampling, and Estimation
by: Fogliato, Riccardo, et al.
Published: (2024) -
Confidence Intervals for Error Rates in 1:1 Matching Tasks: Critical Statistical Analysis and Recommendations
by: Fogliato, Riccardo, et al.
Published: (2023) -
Uncertainty Guarantees on Automated Precision Weeding using Conformal Prediction
by: Melki, Paul, et al.
Published: (2025) -
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
by: Rajabi, Navid, et al.
Published: (2024) -
Classification of Buried Objects from Ground Penetrating Radar Images by using Second Order Deep Learning Models
by: Jafuno, Douba, et al.
Published: (2024)