tinyBenchmarks: evaluating LLMs with fewer examples
Fuente:
arXiv
Salvato in:
| Autori principali: | Polo, Felipe Maia, Weber, Lucas, Choshen, Leshem, Sun, Yuekai, Xu, Gongjun, Yurochkin, Mikhail |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient multi-prompt evaluation of LLMs
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024)
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
di: Polo, Felipe Maia, et al.
Pubblicazione: (2025)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2025)
Weak Supervision Performance Evaluation via Partial Identification
di: Polo, Felipe Maia, et al.
Pubblicazione: (2023)
di: Polo, Felipe Maia, et al.
Pubblicazione: (2023)
Aligners: Decoupling LLMs and Alignment
di: Ngweta, Lilian, et al.
Pubblicazione: (2024)
di: Ngweta, Lilian, et al.
Pubblicazione: (2024)
A Latent Variable Framework for Scaling Laws in Large Language Models
di: Cai, Peiyao, et al.
Pubblicazione: (2025)
di: Cai, Peiyao, et al.
Pubblicazione: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
di: Zaman, Kerem, et al.
Pubblicazione: (2023)
di: Zaman, Kerem, et al.
Pubblicazione: (2023)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
di: Brüel-Gabrielsson, Rickard, et al.
Pubblicazione: (2024)
di: Brüel-Gabrielsson, Rickard, et al.
Pubblicazione: (2024)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
Prompt Exploration with Prompt Regression
di: Feffer, Michael, et al.
Pubblicazione: (2024)
di: Feffer, Michael, et al.
Pubblicazione: (2024)
Robustness as an Emergent Property of Task Performance
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026)
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026)
Out-of-Distribution Detection using Synthetic Data Generation
di: Abbas, Momin, et al.
Pubblicazione: (2025)
di: Abbas, Momin, et al.
Pubblicazione: (2025)
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
di: Yadav, Prateek, et al.
Pubblicazione: (2024)
di: Yadav, Prateek, et al.
Pubblicazione: (2024)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
di: Damani, Mehul, et al.
Pubblicazione: (2025)
di: Damani, Mehul, et al.
Pubblicazione: (2025)
TextArena
di: Guertler, Leon, et al.
Pubblicazione: (2025)
di: Guertler, Leon, et al.
Pubblicazione: (2025)
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
Resolving Interference (RI): Disentangling Models for Improved Model Merging
di: Ramesh, Pratik, et al.
Pubblicazione: (2026)
di: Ramesh, Pratik, et al.
Pubblicazione: (2026)
Fusing Models with Complementary Expertise
di: Wang, Hongyi, et al.
Pubblicazione: (2023)
di: Wang, Hongyi, et al.
Pubblicazione: (2023)
Efficient Benchmarking of Language Models
di: Perlitz, Yotam, et al.
Pubblicazione: (2023)
di: Perlitz, Yotam, et al.
Pubblicazione: (2023)
A transfer learning framework for weak-to-strong generalization
di: Somerstep, Seamus, et al.
Pubblicazione: (2024)
di: Somerstep, Seamus, et al.
Pubblicazione: (2024)
Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT
di: Zain, Noor Ul, et al.
Pubblicazione: (2025)
di: Zain, Noor Ul, et al.
Pubblicazione: (2025)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
di: Ifergan, Maxim, et al.
Pubblicazione: (2024)
di: Ifergan, Maxim, et al.
Pubblicazione: (2024)
Do LLMs Benefit From Their Own Words?
di: Huang, Jenny Y., et al.
Pubblicazione: (2026)
di: Huang, Jenny Y., et al.
Pubblicazione: (2026)
Can Gradient Descent Simulate Prompting?
di: Zhang, Eric, et al.
Pubblicazione: (2025)
di: Zhang, Eric, et al.
Pubblicazione: (2025)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
di: Janiak, Denis, et al.
Pubblicazione: (2025)
di: Janiak, Denis, et al.
Pubblicazione: (2025)
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
di: Murthy, Rithesh, et al.
Pubblicazione: (2024)
di: Murthy, Rithesh, et al.
Pubblicazione: (2024)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
di: Davoodi, Arash Gholami, et al.
Pubblicazione: (2024)
di: Davoodi, Arash Gholami, et al.
Pubblicazione: (2024)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
di: Lin, Zicheng, et al.
Pubblicazione: (2024)
di: Lin, Zicheng, et al.
Pubblicazione: (2024)
Automated test generation to evaluate tool-augmented LLMs as conversational AI agents
di: Arcadinho, Samuel, et al.
Pubblicazione: (2024)
di: Arcadinho, Samuel, et al.
Pubblicazione: (2024)
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
di: Bergsma, Shane, et al.
Pubblicazione: (2025)
di: Bergsma, Shane, et al.
Pubblicazione: (2025)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
di: Petrov, Ivo, et al.
Pubblicazione: (2025)
di: Petrov, Ivo, et al.
Pubblicazione: (2025)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
di: Liu, Yijun, et al.
Pubblicazione: (2024)
di: Liu, Yijun, et al.
Pubblicazione: (2024)
Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
di: Bandarkar, Lucas, et al.
Pubblicazione: (2026)
di: Bandarkar, Lucas, et al.
Pubblicazione: (2026)
Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
di: Chen, Rubing, et al.
Pubblicazione: (2025)
di: Chen, Rubing, et al.
Pubblicazione: (2025)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
di: Yang, Siwei, et al.
Pubblicazione: (2024)
di: Yang, Siwei, et al.
Pubblicazione: (2024)
Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
di: Xiao, Feng, et al.
Pubblicazione: (2025)
di: Xiao, Feng, et al.
Pubblicazione: (2025)
Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models
di: Ivanova, Anna A., et al.
Pubblicazione: (2024)
di: Ivanova, Anna A., et al.
Pubblicazione: (2024)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
di: Ramezanali, Mohammad, et al.
Pubblicazione: (2025)
di: Ramezanali, Mohammad, et al.
Pubblicazione: (2025)
Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework
di: Guan, Zihan, et al.
Pubblicazione: (2026)
di: Guan, Zihan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Efficient multi-prompt evaluation of LLMs
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024) -
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
di: Polo, Felipe Maia, et al.
Pubblicazione: (2024) -
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
di: Polo, Felipe Maia, et al.
Pubblicazione: (2025) -
Weak Supervision Performance Evaluation via Partial Identification
di: Polo, Felipe Maia, et al.
Pubblicazione: (2023) -
Aligners: Decoupling LLMs and Alignment
di: Ngweta, Lilian, et al.
Pubblicazione: (2024)