How predictable is language model benchmark performance?
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Owen, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Comprehensive benchmarking of large language models for RNA secondary structure prediction
von: Zablocki, L. I., et al.
Veröffentlicht: (2024)
von: Zablocki, L. I., et al.
Veröffentlicht: (2024)
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
A dataset and benchmark for hospital course summarization with adapted large language models
von: Aali, Asad, et al.
Veröffentlicht: (2024)
von: Aali, Asad, et al.
Veröffentlicht: (2024)
How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
von: Singh, Prabhant, et al.
Veröffentlicht: (2025)
von: Singh, Prabhant, et al.
Veröffentlicht: (2025)
FlexMS is a flexible framework for benchmarking deep learning-based mass spectrum prediction tools in metabolomics
von: Zhong, Yunhua, et al.
Veröffentlicht: (2026)
von: Zhong, Yunhua, et al.
Veröffentlicht: (2026)
What drives performance in molecular MPNNs? An operator-level factorial benchmark
von: Jiao, Panyu, et al.
Veröffentlicht: (2026)
von: Jiao, Panyu, et al.
Veröffentlicht: (2026)
Integrating remote sensing data assimilation, deep learning and large language model for interactive wheat breeding yield prediction
von: Yang, Guofeng, et al.
Veröffentlicht: (2025)
von: Yang, Guofeng, et al.
Veröffentlicht: (2025)
EXACT: Towards a platform for empirically benchmarking Machine Learning model explanation methods
von: Clark, Benedict, et al.
Veröffentlicht: (2024)
von: Clark, Benedict, et al.
Veröffentlicht: (2024)
Towards impactful challenges: post-challenge paper, benchmarks and other dissemination actions
von: Marot, Antoine, et al.
Veröffentlicht: (2023)
von: Marot, Antoine, et al.
Veröffentlicht: (2023)
Comparative performance of ensemble models in predicting dental provider types: insights from fee-for-service data
von: Al-Batah, Mohammad Subhi, et al.
Veröffentlicht: (2025)
von: Al-Batah, Mohammad Subhi, et al.
Veröffentlicht: (2025)
The role of positional encodings in the ARC benchmark
von: Costa, Guilherme H. Bandeira, et al.
Veröffentlicht: (2025)
von: Costa, Guilherme H. Bandeira, et al.
Veröffentlicht: (2025)
FairX: A comprehensive benchmarking tool for model analysis using fairness, utility, and explainability
von: Sikder, Md Fahim, et al.
Veröffentlicht: (2024)
von: Sikder, Md Fahim, et al.
Veröffentlicht: (2024)
Failure to Mix: Large language models struggle to answer according to desired probability distributions
von: Yang, Ivy Yuqian, et al.
Veröffentlicht: (2025)
von: Yang, Ivy Yuqian, et al.
Veröffentlicht: (2025)
Large language models can accurately predict searcher preferences
von: Thomas, Paul, et al.
Veröffentlicht: (2023)
von: Thomas, Paul, et al.
Veröffentlicht: (2023)
Implicit meta-learning may lead language models to trust more reliable sources
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2023)
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2023)
CXMArena: Unified Dataset to benchmark performance in realistic CXM Scenarios
von: Garg, Raghav, et al.
Veröffentlicht: (2025)
von: Garg, Raghav, et al.
Veröffentlicht: (2025)
Efficiency optimization of large-scale language models based on deep learning in natural language processing tasks
von: Mei, Taiyuan, et al.
Veröffentlicht: (2024)
von: Mei, Taiyuan, et al.
Veröffentlicht: (2024)
PSBench: a large-scale benchmark for estimating the accuracy of protein complex structural models
von: Neupane, Pawan, et al.
Veröffentlicht: (2025)
von: Neupane, Pawan, et al.
Veröffentlicht: (2025)
Applying sparse autoencoders to unlearn knowledge in language models
von: Farrell, Eoin, et al.
Veröffentlicht: (2024)
von: Farrell, Eoin, et al.
Veröffentlicht: (2024)
Quantifying construct validity in large language model evaluations
von: Kearns, Ryan Othniel
Veröffentlicht: (2026)
von: Kearns, Ryan Othniel
Veröffentlicht: (2026)
The OPS-SAT benchmark for detecting anomalies in satellite telemetry
von: Ruszczak, Bogdan, et al.
Veröffentlicht: (2024)
von: Ruszczak, Bogdan, et al.
Veröffentlicht: (2024)
Robust NAS under adversarial training: benchmark, theory, and beyond
von: Wu, Yongtao, et al.
Veröffentlicht: (2024)
von: Wu, Yongtao, et al.
Veröffentlicht: (2024)
A method to benchmark high-dimensional process drift detection
von: Wolf, Edgar, et al.
Veröffentlicht: (2024)
von: Wolf, Edgar, et al.
Veröffentlicht: (2024)
Cueless EEG imagined speech for subject identification: dataset and benchmarks
von: Derakhshesh, Ali, et al.
Veröffentlicht: (2025)
von: Derakhshesh, Ali, et al.
Veröffentlicht: (2025)
Adaptive data selection improves wearable prediction under low baseline performance
von: Kargarandehkordi, Ali
Veröffentlicht: (2026)
von: Kargarandehkordi, Ali
Veröffentlicht: (2026)
The language of time: a language model perspective on time-series foundation models
von: Xie, Yi, et al.
Veröffentlicht: (2025)
von: Xie, Yi, et al.
Veröffentlicht: (2025)
Large language models as uncertainty-calibrated optimizers for experimental discovery
von: Ranković, Bojana, et al.
Veröffentlicht: (2025)
von: Ranković, Bojana, et al.
Veröffentlicht: (2025)
Alignment faking in large language models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
Replacing thinking with tool usage enables reasoning in small language models
von: Rainone, Corrado, et al.
Veröffentlicht: (2025)
von: Rainone, Corrado, et al.
Veröffentlicht: (2025)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
von: Rädsch, Tim, et al.
Veröffentlicht: (2025)
von: Rädsch, Tim, et al.
Veröffentlicht: (2025)
Guaranteed prediction sets for functional surrogate models
von: Gray, Ander, et al.
Veröffentlicht: (2025)
von: Gray, Ander, et al.
Veröffentlicht: (2025)
Vision-language models lag human performance on physical dynamics and intent reasoning
von: Gu, Tianjun, et al.
Veröffentlicht: (2026)
von: Gu, Tianjun, et al.
Veröffentlicht: (2026)
Representation in large language models
von: Yetman, Cameron
Veröffentlicht: (2025)
von: Yetman, Cameron
Veröffentlicht: (2025)
Mixture of A Million Experts
von: He, Xu Owen
Veröffentlicht: (2024)
von: He, Xu Owen
Veröffentlicht: (2024)
Text-guided multi-property molecular optimization with a diffusion language model
von: Xiong, Yida, et al.
Veröffentlicht: (2024)
von: Xiong, Yida, et al.
Veröffentlicht: (2024)
A federated large language model for long-term time series forecasting
von: Abdel-Sater, Raed, et al.
Veröffentlicht: (2024)
von: Abdel-Sater, Raed, et al.
Veröffentlicht: (2024)
Pretraining large language models with MXFP4 on Native FP4 Hardware
von: Cim, Musa, et al.
Veröffentlicht: (2026)
von: Cim, Musa, et al.
Veröffentlicht: (2026)
Diffusion on language model encodings for protein sequence generation
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2024)
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2024)
Assessing win strength in MLB win prediction models
von: Allen, Morgan, et al.
Veröffentlicht: (2025)
von: Allen, Morgan, et al.
Veröffentlicht: (2025)
Hidden markov model to predict tourists visited place
von: Demessance, Theo, et al.
Veröffentlicht: (2025)
von: Demessance, Theo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Comprehensive benchmarking of large language models for RNA secondary structure prediction
von: Zablocki, L. I., et al.
Veröffentlicht: (2024) -
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024) -
A dataset and benchmark for hospital course summarization with adapted large language models
von: Aali, Asad, et al.
Veröffentlicht: (2024) -
How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
von: Singh, Prabhant, et al.
Veröffentlicht: (2025) -
FlexMS is a flexible framework for benchmarking deep learning-based mass spectrum prediction tools in metabolomics
von: Zhong, Yunhua, et al.
Veröffentlicht: (2026)