Cost-Efficient Estimation of General Abilities Across Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Krumdick, Michael, Wiemerslage, Adam, Ebner, Seth, Lovering, Charles, Tanner, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
by: Krumdick, Michael, et al.
Published: (2025)
by: Krumdick, Michael, et al.
Published: (2025)
Tokenization with Split Trees
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
On Finding Inconsistencies in Documents
by: Lovering, Charles J., et al.
Published: (2025)
by: Lovering, Charles J., et al.
Published: (2025)
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
The Effect of Scripts and Formats on LLM Numeracy
by: Reddy, Varshini, et al.
Published: (2026)
by: Reddy, Varshini, et al.
Published: (2026)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024)
by: Lai, Viet Dac, et al.
Published: (2024)
DocFinQA: A Long-Context Financial Reasoning Dataset
by: Reddy, Varshini, et al.
Published: (2024)
by: Reddy, Varshini, et al.
Published: (2024)
From Algebraic Word Problem to Program: A Formalized Approach
by: Wiemerslage, Adam, et al.
Published: (2020)
by: Wiemerslage, Adam, et al.
Published: (2020)
Language Model Probabilities are Not Calibrated in Numeric Contexts
by: Lovering, Charles, et al.
Published: (2024)
by: Lovering, Charles, et al.
Published: (2024)
Improving Low-Resource Morphological Inflection via Self-Supervised Objectives
by: Wiemerslage, Adam, et al.
Published: (2025)
by: Wiemerslage, Adam, et al.
Published: (2025)
Not How Many, But Which: Parameter Placement in Low-Rank Adaptation
by: Sehanobish, Arijit, et al.
Published: (2026)
by: Sehanobish, Arijit, et al.
Published: (2026)
Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual Transfer
by: Ebrahimi, Abteen, et al.
Published: (2025)
by: Ebrahimi, Abteen, et al.
Published: (2025)
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
by: Krumdick, Michael, et al.
Published: (2026)
by: Krumdick, Michael, et al.
Published: (2026)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
by: Chang, Yapei, et al.
Published: (2025)
by: Chang, Yapei, et al.
Published: (2025)
A Closer Look at Claim Decomposition
by: Wanner, Miriam, et al.
Published: (2024)
by: Wanner, Miriam, et al.
Published: (2024)
An Analysis of Multilingual FActScore
by: Vu, Kim Trong, et al.
Published: (2024)
by: Vu, Kim Trong, et al.
Published: (2024)
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
by: Jacovi, Alon, et al.
Published: (2025)
by: Jacovi, Alon, et al.
Published: (2025)
Faster Superword Tokenization
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
Improve LLM-as-a-Judge Ability as a General Ability
by: Yu, Jiachen, et al.
Published: (2025)
by: Yu, Jiachen, et al.
Published: (2025)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
by: Liu, Yijun, et al.
Published: (2024)
by: Liu, Yijun, et al.
Published: (2024)
Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively
by: Gu, Jiawei, et al.
Published: (2025)
by: Gu, Jiawei, et al.
Published: (2025)
Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
by: Ma, Qingchuan, et al.
Published: (2025)
by: Ma, Qingchuan, et al.
Published: (2025)
Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
by: Jahan, Israt, et al.
Published: (2025)
by: Jahan, Israt, et al.
Published: (2025)
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
by: Luo, Zhimeng, et al.
Published: (2025)
by: Luo, Zhimeng, et al.
Published: (2025)
Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries
by: Ide, Yusuke, et al.
Published: (2026)
by: Ide, Yusuke, et al.
Published: (2026)
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
Persuasive or Neutral? A Field Experiment on Generative AI in Online Travel Planning
by: Jirpongopas, Lynna, et al.
Published: (2025)
by: Jirpongopas, Lynna, et al.
Published: (2025)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
by: Uzan, Omri, et al.
Published: (2024)
by: Uzan, Omri, et al.
Published: (2024)
ScholarSearch: Benchmarking Scholar Searching Ability of LLMs
by: Zhou, Junting, et al.
Published: (2025)
by: Zhou, Junting, et al.
Published: (2025)
QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies
by: Khoroshilov, Alexey, et al.
Published: (2026)
by: Khoroshilov, Alexey, et al.
Published: (2026)
M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark
by: Song, Wei, et al.
Published: (2024)
by: Song, Wei, et al.
Published: (2024)
RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation
by: Diallo, Aissatou, et al.
Published: (2025)
by: Diallo, Aissatou, et al.
Published: (2025)
The SIFo Benchmark: Investigating the Sequential Instruction Following Ability of Large Language Models
by: Chen, Xinyi, et al.
Published: (2024)
by: Chen, Xinyi, et al.
Published: (2024)
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
Adaptive Cost-Efficient Evaluation for Reliable Patent Claim Generation
by: Yoo, Yongmin, et al.
Published: (2026)
by: Yoo, Yongmin, et al.
Published: (2026)
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
by: Hagendorff, Thilo, et al.
Published: (2025)
by: Hagendorff, Thilo, et al.
Published: (2025)
BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
by: Zhou, Peilin, et al.
Published: (2025)
by: Zhou, Peilin, et al.
Published: (2025)
Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties
by: Wang, Zhenglin, et al.
Published: (2025)
by: Wang, Zhenglin, et al.
Published: (2025)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
by: Reddy, Varshini, et al.
Published: (2025)
by: Reddy, Varshini, et al.
Published: (2025)
Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network Services
by: Guo, Hongcheng, et al.
Published: (2025)
by: Guo, Hongcheng, et al.
Published: (2025)
Similar Items
-
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
by: Krumdick, Michael, et al.
Published: (2025) -
Tokenization with Split Trees
by: Schmidt, Craig W., et al.
Published: (2026) -
On Finding Inconsistencies in Documents
by: Lovering, Charles J., et al.
Published: (2025) -
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023) -
The Effect of Scripts and Formats on LLM Numeracy
by: Reddy, Varshini, et al.
Published: (2026)