Efficient Benchmarking of Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Perlitz, Yotam, Bandel, Elron, Gera, Ariel, Arviv, Ofir, Ein-Dor, Liat, Shnarch, Eyal, Slonim, Noam, Shmueli-Scheuer, Michal, Choshen, Leshem |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)
by: Perlitz, Yotam, et al.
Published: (2024)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
by: Ashury-Tahan, Shir, et al.
Published: (2025)
by: Ashury-Tahan, Shir, et al.
Published: (2025)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
by: Habba, Eliya, et al.
Published: (2026)
by: Habba, Eliya, et al.
Published: (2026)
Robustness as an Emergent Property of Task Performance
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
Label-Efficient Model Selection for Text Generation
by: Ashury-Tahan, Shir, et al.
Published: (2024)
by: Ashury-Tahan, Shir, et al.
Published: (2024)
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
by: Bandel, Elron, et al.
Published: (2024)
by: Bandel, Elron, et al.
Published: (2024)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
by: Halfon, Alon, et al.
Published: (2024)
by: Halfon, Alon, et al.
Published: (2024)
Instructions Shape Production of Language, not Processing
by: Waldis, Andreas, et al.
Published: (2026)
by: Waldis, Andreas, et al.
Published: (2026)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
by: Yehudai, Asaf, et al.
Published: (2024)
by: Yehudai, Asaf, et al.
Published: (2024)
WildIFEval: Instruction Following in the Wild
by: Lior, Gili, et al.
Published: (2025)
by: Lior, Gili, et al.
Published: (2025)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
by: Uzan, Omri, et al.
Published: (2025)
by: Uzan, Omri, et al.
Published: (2025)
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
by: Sternlicht, Noy, et al.
Published: (2025)
by: Sternlicht, Noy, et al.
Published: (2025)
General Agent Evaluation
by: Bandel, Elron, et al.
Published: (2026)
by: Bandel, Elron, et al.
Published: (2026)
Conversational Prompt Engineering
by: Ein-Dor, Liat, et al.
Published: (2024)
by: Ein-Dor, Liat, et al.
Published: (2024)
Multi-Domain Explainability of Preferences
by: Calderon, Nitay, et al.
Published: (2025)
by: Calderon, Nitay, et al.
Published: (2025)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs
by: Peisakhovsky, Yehonatan, et al.
Published: (2025)
by: Peisakhovsky, Yehonatan, et al.
Published: (2025)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
by: Ashuach, Tomer, et al.
Published: (2026)
by: Ashuach, Tomer, et al.
Published: (2026)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
by: Kour, George, et al.
Published: (2025)
by: Kour, George, et al.
Published: (2025)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Naturally Occurring Feedback is Common, Extractable and Useful
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Can Gradient Descent Simulate Prompting?
by: Zhang, Eric, et al.
Published: (2025)
by: Zhang, Eric, et al.
Published: (2025)
Pretraining Language Models for Diachronic Linguistic Change Discovery
by: Fittschen, Elisabeth, et al.
Published: (2025)
by: Fittschen, Elisabeth, et al.
Published: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
A Hitchhiker's Guide to Scaling Law Estimation
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
by: Zaman, Kerem, et al.
Published: (2023)
by: Zaman, Kerem, et al.
Published: (2023)
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
by: Akyürek, Afra Feyza, et al.
Published: (2024)
by: Akyürek, Afra Feyza, et al.
Published: (2024)
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
by: Keren, Tomer, et al.
Published: (2026)
by: Keren, Tomer, et al.
Published: (2026)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
by: Din, Alexander Yom, et al.
Published: (2023)
by: Din, Alexander Yom, et al.
Published: (2023)
Mediocrity is the key for LLM as a Judge Anchor Selection
by: Don-Yehiya, Shachar, et al.
Published: (2026)
by: Don-Yehiya, Shachar, et al.
Published: (2026)
Teaching Models to Improve on Tape
by: Bezalel, Liat, et al.
Published: (2024)
by: Bezalel, Liat, et al.
Published: (2024)
PolySQL: Scaling Text-to-SQL Evaluation Across SQL Dialects via Automated Backend Isomorphism
by: Perlitz, Yotam, et al.
Published: (2026)
by: Perlitz, Yotam, et al.
Published: (2026)
NeurIPS 2023 LLM Efficiency Fine-tuning Competition
by: Saroufim, Mark, et al.
Published: (2025)
by: Saroufim, Mark, et al.
Published: (2025)
tinyBenchmarks: evaluating LLMs with fewer examples
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
Similar Items
-
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024) -
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
by: Habba, Eliya, et al.
Published: (2025) -
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
by: Ashury-Tahan, Shir, et al.
Published: (2025) -
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
by: Habba, Eliya, et al.
Published: (2026) -
Robustness as an Emergent Property of Task Performance
by: Ashury-Tahan, Shir, et al.
Published: (2026)