Do Large Language Model Benchmarks Test Reliability?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vendrow, Joshua, Vendrow, Edward, Beery, Sara, Madry, Aleksander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ask Your Distribution Shift if Pre-Training is Right for You
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2024)
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2024)
ContextCite: Attributing Model Generation to Context
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2024)
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2024)
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
Learning to Attribute with Attention
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2025)
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2025)
How Reliable is Language Model Micro-Benchmarking?
von: Yauney, Gregory, et al.
Veröffentlicht: (2025)
von: Yauney, Gregory, et al.
Veröffentlicht: (2025)
On the Reliability of Watermarks for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
Small-to-Large Generalization: Data Influences Models Consistently Across Scale
von: Khaddaj, Alaa, et al.
Veröffentlicht: (2025)
von: Khaddaj, Alaa, et al.
Veröffentlicht: (2025)
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
von: Banerjee, Sourav, et al.
Veröffentlicht: (2024)
von: Banerjee, Sourav, et al.
Veröffentlicht: (2024)
Discovering Implicit Large Language Model Alignment Objectives
von: Chen, Edward, et al.
Veröffentlicht: (2026)
von: Chen, Edward, et al.
Veröffentlicht: (2026)
Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
CEQuest: Benchmarking Large Language Models for Construction Estimation
von: Wu, Yanzhao, et al.
Veröffentlicht: (2025)
von: Wu, Yanzhao, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Model Uncertainty for Prompt Optimization
von: Guo, Pei-Fu, et al.
Veröffentlicht: (2024)
von: Guo, Pei-Fu, et al.
Veröffentlicht: (2024)
Benchmarking Large Language Models for Math Reasoning Tasks
von: Seßler, Kathrin, et al.
Veröffentlicht: (2024)
von: Seßler, Kathrin, et al.
Veröffentlicht: (2024)
Large-Scale, Longitudinal Study of Large Language Models During the 2024 US Election Season
von: Cen, Sarah H., et al.
Veröffentlicht: (2025)
von: Cen, Sarah H., et al.
Veröffentlicht: (2025)
Sequence-Level Leakage Risk of Training Data in Large Language Models
von: Tiwari, Trishita, et al.
Veröffentlicht: (2024)
von: Tiwari, Trishita, et al.
Veröffentlicht: (2024)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
von: Wang, Wentian, et al.
Veröffentlicht: (2024)
von: Wang, Wentian, et al.
Veröffentlicht: (2024)
Benchmarking Benchmark Leakage in Large Language Models
von: Xu, Ruijie, et al.
Veröffentlicht: (2024)
von: Xu, Ruijie, et al.
Veröffentlicht: (2024)
DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models
von: Hu, Xiaolin, et al.
Veröffentlicht: (2024)
von: Hu, Xiaolin, et al.
Veröffentlicht: (2024)
Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models
von: Siddiqui, S M Tahmid, et al.
Veröffentlicht: (2026)
von: Siddiqui, S M Tahmid, et al.
Veröffentlicht: (2026)
Nevermind: Instruction Override and Moderation in Large Language Models
von: Kim, Edward
Veröffentlicht: (2024)
von: Kim, Edward
Veröffentlicht: (2024)
CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
von: Wang, Song, et al.
Veröffentlicht: (2024)
von: Wang, Song, et al.
Veröffentlicht: (2024)
GLBench: A Comprehensive Benchmark for Graph with Large Language Models
von: Li, Yuhan, et al.
Veröffentlicht: (2024)
von: Li, Yuhan, et al.
Veröffentlicht: (2024)
Test-Time Training on Nearest Neighbors for Large Language Models
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
Testing Uncertainty of Large Language Models for Physics Knowledge and Reasoning
von: Reganova, Elizaveta, et al.
Veröffentlicht: (2024)
von: Reganova, Elizaveta, et al.
Veröffentlicht: (2024)
Objective Metrics for Evaluating Large Language Models Using External Data Sources
von: Du, Haoze, et al.
Veröffentlicht: (2025)
von: Du, Haoze, et al.
Veröffentlicht: (2025)
Unmasking Hallucinations: A Causal Graph-Attention Perspective on Factual Reliability in Large Language Models
von: kurra, Sailesh kiran, et al.
Veröffentlicht: (2026)
von: kurra, Sailesh kiran, et al.
Veröffentlicht: (2026)
Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph
von: Vashurin, Roman, et al.
Veröffentlicht: (2024)
von: Vashurin, Roman, et al.
Veröffentlicht: (2024)
metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
von: Kipnis, Alex, et al.
Veröffentlicht: (2024)
von: Kipnis, Alex, et al.
Veröffentlicht: (2024)
A Critical Review of Causal Reasoning Benchmarks for Large Language Models
von: Yang, Linying, et al.
Veröffentlicht: (2024)
von: Yang, Linying, et al.
Veröffentlicht: (2024)
On the Emergence and Test-Time Use of Structural Information in Large Language Models
von: Chen, Michelle Chao, et al.
Veröffentlicht: (2026)
von: Chen, Michelle Chao, et al.
Veröffentlicht: (2026)
Model Provenance Testing for Large Language Models
von: Nikolic, Ivica, et al.
Veröffentlicht: (2025)
von: Nikolic, Ivica, et al.
Veröffentlicht: (2025)
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
von: Yuan, Yu, et al.
Veröffentlicht: (2024)
von: Yuan, Yu, et al.
Veröffentlicht: (2024)
TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
von: Li, Jianling, et al.
Veröffentlicht: (2025)
von: Li, Jianling, et al.
Veröffentlicht: (2025)
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
Benchmarking Uncertainty Calibration in Large Language Model Long-Form Question Answering
von: Müller, Philip, et al.
Veröffentlicht: (2026)
von: Müller, Philip, et al.
Veröffentlicht: (2026)
DetoxBench: Benchmarking Large Language Models for Multitask Fraud & Abuse Detection
von: Chakraborty, Joymallya, et al.
Veröffentlicht: (2024)
von: Chakraborty, Joymallya, et al.
Veröffentlicht: (2024)
Do Large Language Models Need Intent? Revisiting Response Generation Strategies for Service Assistant
von: Bolshinsky, Inbal, et al.
Veröffentlicht: (2025)
von: Bolshinsky, Inbal, et al.
Veröffentlicht: (2025)
Test-Time Learning for Large Language Models
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models
von: Mateega, Spencer, et al.
Veröffentlicht: (2025)
von: Mateega, Spencer, et al.
Veröffentlicht: (2025)
MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ask Your Distribution Shift if Pre-Training is Right for You
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2024) -
ContextCite: Attributing Model Generation to Context
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2024) -
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
von: Vendrow, Edward, et al.
Veröffentlicht: (2024) -
Learning to Attribute with Attention
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2025) -
How Reliable is Language Model Micro-Benchmarking?
von: Yauney, Gregory, et al.
Veröffentlicht: (2025)