Efficient multi-prompt evaluation of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Polo, Felipe Maia, Xu, Ronald, Weber, Lucas, Silva, Mírian, Bhardwaj, Onkar, Choshen, Leshem, de Oliveira, Allysson Flavio Melo, Sun, Yuekai, Yurochkin, Mikhail |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
tinyBenchmarks: evaluating LLMs with fewer examples
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
CARROT: A Cost Aware Rate Optimal Router
by: Somerstep, Seamus, et al.
Published: (2025)
by: Somerstep, Seamus, et al.
Published: (2025)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
Weak Supervision Performance Evaluation via Partial Identification
by: Polo, Felipe Maia, et al.
Published: (2023)
by: Polo, Felipe Maia, et al.
Published: (2023)
A Latent Variable Framework for Scaling Laws in Large Language Models
by: Cai, Peiyao, et al.
Published: (2025)
by: Cai, Peiyao, et al.
Published: (2025)
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
by: Polo, Felipe Maia, et al.
Published: (2025)
by: Polo, Felipe Maia, et al.
Published: (2025)
Fusing Models with Complementary Expertise
by: Wang, Hongyi, et al.
Published: (2023)
by: Wang, Hongyi, et al.
Published: (2023)
Prompt Exploration with Prompt Regression
by: Feffer, Michael, et al.
Published: (2024)
by: Feffer, Michael, et al.
Published: (2024)
A transfer learning framework for weak-to-strong generalization
by: Somerstep, Seamus, et al.
Published: (2024)
by: Somerstep, Seamus, et al.
Published: (2024)
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
by: Shabtay, Nimrod, et al.
Published: (2024)
by: Shabtay, Nimrod, et al.
Published: (2024)
Aligners: Decoupling LLMs and Alignment
by: Ngweta, Lilian, et al.
Published: (2024)
by: Ngweta, Lilian, et al.
Published: (2024)
Limitations of refinement methods for weak to strong generalization
by: Somerstep, Seamus, et al.
Published: (2025)
by: Somerstep, Seamus, et al.
Published: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Can Gradient Descent Simulate Prompting?
by: Zhang, Eric, et al.
Published: (2025)
by: Zhang, Eric, et al.
Published: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
by: Zaman, Kerem, et al.
Published: (2023)
by: Zaman, Kerem, et al.
Published: (2023)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
Do LLMs Benefit From Their Own Words?
by: Huang, Jenny Y., et al.
Published: (2026)
by: Huang, Jenny Y., et al.
Published: (2026)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Naturally Occurring Feedback is Common, Extractable and Useful
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Microfoundation Inference for Strategic Prediction
by: Bracale, Daniele, et al.
Published: (2024)
by: Bracale, Daniele, et al.
Published: (2024)
Asymmetry in Low-Rank Adapters of Foundation Models
by: Zhu, Jiacheng, et al.
Published: (2024)
by: Zhu, Jiacheng, et al.
Published: (2024)
NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning
by: Schwartz, Eli, et al.
Published: (2024)
by: Schwartz, Eli, et al.
Published: (2024)
Instructions Shape Production of Language, not Processing
by: Waldis, Andreas, et al.
Published: (2026)
by: Waldis, Andreas, et al.
Published: (2026)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
by: Din, Alexander Yom, et al.
Published: (2023)
by: Din, Alexander Yom, et al.
Published: (2023)
Mediocrity is the key for LLM as a Judge Anchor Selection
by: Don-Yehiya, Shachar, et al.
Published: (2026)
by: Don-Yehiya, Shachar, et al.
Published: (2026)
Label-Efficient Model Selection for Text Generation
by: Ashury-Tahan, Shir, et al.
Published: (2024)
by: Ashury-Tahan, Shir, et al.
Published: (2024)
Pretraining Language Models for Diachronic Linguistic Change Discovery
by: Fittschen, Elisabeth, et al.
Published: (2025)
by: Fittschen, Elisabeth, et al.
Published: (2025)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
by: Hilel, Almog, et al.
Published: (2025)
by: Hilel, Almog, et al.
Published: (2025)
Resolving Interference (RI): Disentangling Models for Improved Model Merging
by: Ramesh, Pratik, et al.
Published: (2026)
by: Ramesh, Pratik, et al.
Published: (2026)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Model merging with SVD to tie the Knots
by: Stoica, George, et al.
Published: (2024)
by: Stoica, George, et al.
Published: (2024)
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
by: Akyürek, Afra Feyza, et al.
Published: (2024)
by: Akyürek, Afra Feyza, et al.
Published: (2024)
Covid-19 in Brazil: stress as predictor of depression
by: Silva, Washington Allysson Dantas
Published: (2020)
by: Silva, Washington Allysson Dantas
Published: (2020)
A Cataguases de Luiz Ruffato: a experiência urbana no interior mineiro
by: Allysson Augusto Silva Casais
Published: (2022)
by: Allysson Augusto Silva Casais
Published: (2022)
A marginalização ocupa a rua: O rap em Cuba
by: Allysson Fernandes
Published: (2014)
by: Allysson Fernandes
Published: (2014)
Will it Merge? On The Causes of Model Mergeability
by: Rahamim, Adir, et al.
Published: (2026)
by: Rahamim, Adir, et al.
Published: (2026)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
TextArena
by: Guertler, Leon, et al.
Published: (2025)
by: Guertler, Leon, et al.
Published: (2025)
Similar Items
-
tinyBenchmarks: evaluating LLMs with fewer examples
by: Polo, Felipe Maia, et al.
Published: (2024) -
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
by: Polo, Felipe Maia, et al.
Published: (2024) -
CARROT: A Cost Aware Rate Optimal Router
by: Somerstep, Seamus, et al.
Published: (2025) -
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024) -
Weak Supervision Performance Evaluation via Partial Identification
by: Polo, Felipe Maia, et al.
Published: (2023)