CalArena: A Large-Scale Post-Hoc Calibration Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Berta, Eugène, Holzmüller, David, Bach, Francis, Jordan, Michael I. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structured Matrix Scaling for Multi-Class Calibration
by: Berta, Eugène, et al.
Published: (2025)
by: Berta, Eugène, et al.
Published: (2025)
Rethinking Early Stopping: Refine, Then Calibrate
by: Berta, Eugène, et al.
Published: (2025)
by: Berta, Eugène, et al.
Published: (2025)
A Variational Estimator for $L_p$ Calibration Errors
by: Berta, Eugène, et al.
Published: (2026)
by: Berta, Eugène, et al.
Published: (2026)
Conditional Coverage Diagnostics for Conformal Prediction
by: Braun, Sacha, et al.
Published: (2025)
by: Braun, Sacha, et al.
Published: (2025)
Multivariate Standardized Residuals for Conformal Prediction
by: Braun, Sacha, et al.
Published: (2025)
by: Braun, Sacha, et al.
Published: (2025)
TabArena: A Living Benchmark for Machine Learning on Tabular Data
by: Erickson, Nick, et al.
Published: (2025)
by: Erickson, Nick, et al.
Published: (2025)
FeatCal: Feature Calibration for Post-Merging Models
by: Gu, Yanggan, et al.
Published: (2026)
by: Gu, Yanggan, et al.
Published: (2026)
Super-Level-Set Regression: Conditional Quantiles via Volume Minimization
by: Braun, Sacha, et al.
Published: (2026)
by: Braun, Sacha, et al.
Published: (2026)
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
by: Qu, Jingang, et al.
Published: (2025)
by: Qu, Jingang, et al.
Published: (2025)
Convergence Rates for Non-Log-Concave Sampling and Log-Partition Estimation
by: Holzmüller, David, et al.
Published: (2023)
by: Holzmüller, David, et al.
Published: (2023)
Position: Rethinking Post-Hoc Search-Based Neural Approaches for Solving Large-Scale Traveling Salesman Problems
by: Xia, Yifan, et al.
Published: (2024)
by: Xia, Yifan, et al.
Published: (2024)
Post Hoc Regression Refinement via Pairwise Rankings
by: Wijaya, Kevin Tirta, et al.
Published: (2025)
by: Wijaya, Kevin Tirta, et al.
Published: (2025)
ChaosMining: A Benchmark to Evaluate Post-Hoc Local Attribution Methods in Low SNR Environments
by: Shi, Ge, et al.
Published: (2024)
by: Shi, Ge, et al.
Published: (2024)
Scalable Utility-Aware Multiclass Calibration
by: Hegazy, Mahmoud, et al.
Published: (2025)
by: Hegazy, Mahmoud, et al.
Published: (2025)
Minimum Volume Conformal Sets for Multivariate Regression
by: Braun, Sacha, et al.
Published: (2025)
by: Braun, Sacha, et al.
Published: (2025)
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
Post-Hoc Reversal: Are We Selecting Models Prematurely?
by: Ranjan, Rishabh, et al.
Published: (2024)
by: Ranjan, Rishabh, et al.
Published: (2024)
Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
by: Nakamoto, Mitsuhiko, et al.
Published: (2023)
by: Nakamoto, Mitsuhiko, et al.
Published: (2023)
Comment on paper: Position: Rethinking Post-Hoc Search-Based Neural Approaches for Solving Large-Scale Traveling Salesman Problems
by: Min, Yimeng
Published: (2024)
by: Min, Yimeng
Published: (2024)
CHARM: Calibrating Reward Models With Chatbot Arena Scores
by: Zhu, Xiao, et al.
Published: (2025)
by: Zhu, Xiao, et al.
Published: (2025)
Uncertainty-Aware Post-Hoc Calibration: Mitigating Confidently Incorrect Predictions Beyond Calibration Metrics
by: Gharoun, Hassan, et al.
Published: (2025)
by: Gharoun, Hassan, et al.
Published: (2025)
Agent-Based Post-Hoc Correction of Agricultural Yield Forecasts
by: Beddows, Matthew, et al.
Published: (2026)
by: Beddows, Matthew, et al.
Published: (2026)
Smooth InfoMax -- Towards Easier Post-Hoc Interpretability
by: Denoodt, Fabian, et al.
Published: (2024)
by: Denoodt, Fabian, et al.
Published: (2024)
Informative Post-Hoc Explanations Only Exist for Simple Functions
by: Günther, Eric, et al.
Published: (2025)
by: Günther, Eric, et al.
Published: (2025)
The Galerkin method beats Graph-Based Approaches for Spectral Algorithms
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining
by: Abro, Aarash, et al.
Published: (2026)
by: Abro, Aarash, et al.
Published: (2026)
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
by: Yan, Xianglong, et al.
Published: (2025)
by: Yan, Xianglong, et al.
Published: (2025)
Operationalizing Fairness: Post-Hoc Threshold Optimization Under Hard Resource Limits
by: Singh, Moirangthem Tiken, et al.
Published: (2026)
by: Singh, Moirangthem Tiken, et al.
Published: (2026)
Principled Input-Output-Conditioned Post-Hoc Uncertainty Estimation for Regression Networks
by: Bramlage, Lennart, et al.
Published: (2025)
by: Bramlage, Lennart, et al.
Published: (2025)
Are We Merely Justifying Results ex Post Facto? Quantifying Explanatory Inversion in Post-Hoc Model Explanations
by: Tan, Zhen, et al.
Published: (2025)
by: Tan, Zhen, et al.
Published: (2025)
NetArena: Dynamic Benchmarks for AI Agents in Network Automation
by: Zhou, Yajie, et al.
Published: (2025)
by: Zhou, Yajie, et al.
Published: (2025)
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models
by: Wang, Jason Z
Published: (2026)
by: Wang, Jason Z
Published: (2026)
ShapeX: Shapelet-Driven Post Hoc Explanations for Time Series Classification Models
by: Huang, Bosong, et al.
Published: (2025)
by: Huang, Bosong, et al.
Published: (2025)
A Benchmark Study on Calibration
by: Tao, Linwei, et al.
Published: (2023)
by: Tao, Linwei, et al.
Published: (2023)
TaylorPODA: A Taylor Expansion-Based Method to Improve Post-Hoc Attributions for Opaque Models
by: Tang, Yuchi, et al.
Published: (2025)
by: Tang, Yuchi, et al.
Published: (2025)
LOGLO-FNO: Efficient Learning of Local and Global Features in Fourier Neural Operators
by: Kalimuthu, Marimuthu, et al.
Published: (2025)
by: Kalimuthu, Marimuthu, et al.
Published: (2025)
SHapley Estimated Explanation (SHEP): A Fast Post-Hoc Attribution Method for Interpreting Intelligent Fault Diagnosis
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
by: Luo, Haipeng, et al.
Published: (2024)
by: Luo, Haipeng, et al.
Published: (2024)
xai_evals : A Framework for Evaluating Post-Hoc Local Explanation Methods
by: Seth, Pratinav, et al.
Published: (2025)
by: Seth, Pratinav, et al.
Published: (2025)
Comparing Post-Hoc Explainable AI Methods for Interpreting Black-Box EEG Models in Depression Detection
by: Šarčević, Antonia, et al.
Published: (2026)
by: Šarčević, Antonia, et al.
Published: (2026)
Similar Items
-
Structured Matrix Scaling for Multi-Class Calibration
by: Berta, Eugène, et al.
Published: (2025) -
Rethinking Early Stopping: Refine, Then Calibrate
by: Berta, Eugène, et al.
Published: (2025) -
A Variational Estimator for $L_p$ Calibration Errors
by: Berta, Eugène, et al.
Published: (2026) -
Conditional Coverage Diagnostics for Conformal Prediction
by: Braun, Sacha, et al.
Published: (2025) -
Multivariate Standardized Residuals for Conformal Prediction
by: Braun, Sacha, et al.
Published: (2025)