The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Banerjee, Sourav, Agarwal, Ayushi, Singh, Eishkaran |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI
par: Banerjee, Sourav, et autres
Publié: (2024)
par: Banerjee, Sourav, et autres
Publié: (2024)
LLMs Will Always Hallucinate, and We Need to Live With This
par: Banerjee, Sourav, et autres
Publié: (2024)
par: Banerjee, Sourav, et autres
Publié: (2024)
High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
par: Banerjee, Sourav, et autres
Publié: (2024)
par: Banerjee, Sourav, et autres
Publié: (2024)
Do Large Language Model Benchmarks Test Reliability?
par: Vendrow, Joshua, et autres
Publié: (2025)
par: Vendrow, Joshua, et autres
Publié: (2025)
Narrative Analysis of True Crime Podcasts With Knowledge Graph-Augmented Large Language Models
par: Leng, Xinyi, et autres
Publié: (2024)
par: Leng, Xinyi, et autres
Publié: (2024)
Causal Reflection with Language Models
par: Aryan, Abi, et autres
Publié: (2025)
par: Aryan, Abi, et autres
Publié: (2025)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
par: Nirmal, Ayushi, et autres
Publié: (2024)
par: Nirmal, Ayushi, et autres
Publié: (2024)
nanoLM: an Affordable LLM Pre-training Benchmark via Accurate Loss Prediction across Scales
par: Yao, Yiqun, et autres
Publié: (2023)
par: Yao, Yiqun, et autres
Publié: (2023)
Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings
par: Wani, Sachin Gopal, et autres
Publié: (2026)
par: Wani, Sachin Gopal, et autres
Publié: (2026)
PATENTWRITER: A Benchmarking Study for Patent Drafting with LLMs
par: Shomee, Homaira Huda, et autres
Publié: (2025)
par: Shomee, Homaira Huda, et autres
Publié: (2025)
Leviathan: Decoupling Input and Output Representations in Language Models
par: Batley, Reza T., et autres
Publié: (2026)
par: Batley, Reza T., et autres
Publié: (2026)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
par: Schaeffer, Rylan, et autres
Publié: (2025)
par: Schaeffer, Rylan, et autres
Publié: (2025)
Large Language Models Reflect the Ideology of their Creators
par: Buyl, Maarten, et autres
Publié: (2024)
par: Buyl, Maarten, et autres
Publié: (2024)
Representation Consistency for Accurate and Coherent LLM Answer Aggregation
par: Jiang, Junqi, et autres
Publié: (2025)
par: Jiang, Junqi, et autres
Publié: (2025)
Latent Performance Profiling of Large Language Models
par: Chakraborty, Tanmoy, et autres
Publié: (2026)
par: Chakraborty, Tanmoy, et autres
Publié: (2026)
The Illusionist's Prompt: Exposing the Factual Vulnerabilities of Large Language Models with Linguistic Nuances
par: Wang, Yining, et autres
Publié: (2025)
par: Wang, Yining, et autres
Publié: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
par: Kim, Eunsu, et autres
Publié: (2025)
par: Kim, Eunsu, et autres
Publié: (2025)
A Single-Layer Model Can Do Language Modeling
par: Wang, Zanmin
Publié: (2026)
par: Wang, Zanmin
Publié: (2026)
What Evidence Do Language Models Find Convincing?
par: Wan, Alexander, et autres
Publié: (2024)
par: Wan, Alexander, et autres
Publié: (2024)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
par: Kim, Jang-Hyun, et autres
Publié: (2026)
par: Kim, Jang-Hyun, et autres
Publié: (2026)
Aligning Large Language Models to Low-Resource Languages through LLM-Based Selective Translation: A Systematic Study
par: Paul, Rakesh, et autres
Publié: (2025)
par: Paul, Rakesh, et autres
Publié: (2025)
More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents
par: Xiong, Weimin, et autres
Publié: (2025)
par: Xiong, Weimin, et autres
Publié: (2025)
How Reliable is Language Model Micro-Benchmarking?
par: Yauney, Gregory, et autres
Publié: (2025)
par: Yauney, Gregory, et autres
Publié: (2025)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
par: Setlur, Amrith, et autres
Publié: (2024)
par: Setlur, Amrith, et autres
Publié: (2024)
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
par: Zhao, Ziqi, et autres
Publié: (2026)
par: Zhao, Ziqi, et autres
Publié: (2026)
Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models
par: Dai, Runpeng, et autres
Publié: (2025)
par: Dai, Runpeng, et autres
Publié: (2025)
Synergizing Unsupervised and Supervised Learning: A Hybrid Approach for Accurate Natural Language Task Modeling
par: Talukdar, Wrick, et autres
Publié: (2024)
par: Talukdar, Wrick, et autres
Publié: (2024)
ROSE: Reordered SparseGPT for More Accurate One-Shot Large Language Models Pruning
par: Su, Mingluo, et autres
Publié: (2026)
par: Su, Mingluo, et autres
Publié: (2026)
REAMS: Reasoning Enhanced Algorithm for Maths Solving
par: Singh, Eishkaran, et autres
Publié: (2025)
par: Singh, Eishkaran, et autres
Publié: (2025)
PhyloLM : Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks
par: Yax, Nicolas, et autres
Publié: (2024)
par: Yax, Nicolas, et autres
Publié: (2024)
Benchmarking Large Language Model Uncertainty for Prompt Optimization
par: Guo, Pei-Fu, et autres
Publié: (2024)
par: Guo, Pei-Fu, et autres
Publié: (2024)
Benchmarking Large Language Models for Math Reasoning Tasks
par: Seßler, Kathrin, et autres
Publié: (2024)
par: Seßler, Kathrin, et autres
Publié: (2024)
On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
par: Gupta, Aarav, et autres
Publié: (2026)
par: Gupta, Aarav, et autres
Publié: (2026)
CEQuest: Benchmarking Large Language Models for Construction Estimation
par: Wu, Yanzhao, et autres
Publié: (2025)
par: Wu, Yanzhao, et autres
Publié: (2025)
LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
par: Shojaee, Parshin, et autres
Publié: (2025)
par: Shojaee, Parshin, et autres
Publié: (2025)
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
par: Younsi, Adam, et autres
Publié: (2025)
par: Younsi, Adam, et autres
Publié: (2025)
Performance Law of Large Language Models
par: Wu, Chuhan, et autres
Publié: (2024)
par: Wu, Chuhan, et autres
Publié: (2024)
DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models
par: Hu, Xiaolin, et autres
Publié: (2024)
par: Hu, Xiaolin, et autres
Publié: (2024)
What Do Language Models Learn in Context? The Structured Task Hypothesis
par: Li, Jiaoda, et autres
Publié: (2024)
par: Li, Jiaoda, et autres
Publié: (2024)
Discursive Circuits: How Do Language Models Understand Discourse Relations?
par: Miao, Yisong, et autres
Publié: (2025)
par: Miao, Yisong, et autres
Publié: (2025)
Documents similaires
-
First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI
par: Banerjee, Sourav, et autres
Publié: (2024) -
LLMs Will Always Hallucinate, and We Need to Live With This
par: Banerjee, Sourav, et autres
Publié: (2024) -
High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
par: Banerjee, Sourav, et autres
Publié: (2024) -
Do Large Language Model Benchmarks Test Reliability?
par: Vendrow, Joshua, et autres
Publié: (2025) -
Narrative Analysis of True Crime Podcasts With Knowledge Graph-Augmented Large Language Models
par: Leng, Xinyi, et autres
Publié: (2024)