Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ailem, Melissa, Marazopoulou, Katerina, Siska, Charlotte, Bono, James |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
von: Siska, Charlotte, et al.
Veröffentlicht: (2025)
von: Siska, Charlotte, et al.
Veröffentlicht: (2025)
From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
von: Roemmele, Melissa, et al.
Veröffentlicht: (2024)
von: Roemmele, Melissa, et al.
Veröffentlicht: (2024)
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
von: Liu, Shuyu, et al.
Veröffentlicht: (2025)
von: Liu, Shuyu, et al.
Veröffentlicht: (2025)
Efficient Uncertainty Estimation for LLM-based Entity Linking in Tabular Data
von: Bono, Carlo, et al.
Veröffentlicht: (2025)
von: Bono, Carlo, et al.
Veröffentlicht: (2025)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
von: Wen, Zichen, et al.
Veröffentlicht: (2024)
von: Wen, Zichen, et al.
Veröffentlicht: (2024)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
von: Li, Kunning, et al.
Veröffentlicht: (2025)
von: Li, Kunning, et al.
Veröffentlicht: (2025)
CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance
von: Zhang, Feng, et al.
Veröffentlicht: (2025)
von: Zhang, Feng, et al.
Veröffentlicht: (2025)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
von: Montalan, Jann Railey, et al.
Veröffentlicht: (2025)
von: Montalan, Jann Railey, et al.
Veröffentlicht: (2025)
AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2024)
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2024)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
von: Jassim, Serwan, et al.
Veröffentlicht: (2023)
von: Jassim, Serwan, et al.
Veröffentlicht: (2023)
Dynamic benchmarking framework for LLM-based conversational data capture
von: Aluffi, Pietro Alessandro, et al.
Veröffentlicht: (2025)
von: Aluffi, Pietro Alessandro, et al.
Veröffentlicht: (2025)
LLMzSzŁ: a comprehensive LLM benchmark for Polish
von: Jassem, Krzysztof, et al.
Veröffentlicht: (2025)
von: Jassem, Krzysztof, et al.
Veröffentlicht: (2025)
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
von: Kankowski, Florian, et al.
Veröffentlicht: (2025)
von: Kankowski, Florian, et al.
Veröffentlicht: (2025)
OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities
von: Verma, Sahil, et al.
Veröffentlicht: (2025)
von: Verma, Sahil, et al.
Veröffentlicht: (2025)
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
von: Panagoulias, Dimitrios P., et al.
Veröffentlicht: (2024)
von: Panagoulias, Dimitrios P., et al.
Veröffentlicht: (2024)
Examining Identity Drift in Conversations of LLM Agents
von: Choi, Junhyuk, et al.
Veröffentlicht: (2024)
von: Choi, Junhyuk, et al.
Veröffentlicht: (2024)
Verbalizing LLMs' assumptions to explain and control sycophancy
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
Cross-lingual robustness of LLM-brain alignment and its computational roots
von: Yang, Ni, et al.
Veröffentlicht: (2026)
von: Yang, Ni, et al.
Veröffentlicht: (2026)
Monte Carlo Temperature: a robust sampling strategy for LLM's uncertainty quantification methods
von: Cecere, Nicola, et al.
Veröffentlicht: (2025)
von: Cecere, Nicola, et al.
Veröffentlicht: (2025)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2024)
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2024)
Examining the Behavior of LLM Architectures Within the Framework of Standardized National Exams in Brazil
von: Locatelli, Marcelo Sartori, et al.
Veröffentlicht: (2024)
von: Locatelli, Marcelo Sartori, et al.
Veröffentlicht: (2024)
Linguini: A benchmark for language-agnostic linguistic reasoning
von: Sánchez, Eduardo, et al.
Veröffentlicht: (2024)
von: Sánchez, Eduardo, et al.
Veröffentlicht: (2024)
An Expert-grounded benchmark of General Purpose LLMs in LCA
von: Donaldson, Artur, et al.
Veröffentlicht: (2025)
von: Donaldson, Artur, et al.
Veröffentlicht: (2025)
How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives
von: Ichmoukhamedov, Timour, et al.
Veröffentlicht: (2024)
von: Ichmoukhamedov, Timour, et al.
Veröffentlicht: (2024)
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples
von: Michail, Andrianos, et al.
Veröffentlicht: (2025)
von: Michail, Andrianos, et al.
Veröffentlicht: (2025)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
Certainty robustness: Evaluating LLM stability under self-challenging prompts
von: Saadat, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Saadat, Mohammadreza, et al.
Veröffentlicht: (2026)
GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
von: Bandooni, Ashutosh, et al.
Veröffentlicht: (2025)
von: Bandooni, Ashutosh, et al.
Veröffentlicht: (2025)
NLP for The Greek Language: A Longer Survey
von: Papantoniou, Katerina, et al.
Veröffentlicht: (2024)
von: Papantoniou, Katerina, et al.
Veröffentlicht: (2024)
Proverbs or Pythian Oracles? Sentiments and Emotions in Greek Sayings
von: Korre, Katerina, et al.
Veröffentlicht: (2025)
von: Korre, Katerina, et al.
Veröffentlicht: (2025)
BIPOLAR: Polarization-based granular framework for LLM bias evaluation
von: Pavlíček, Martin, et al.
Veröffentlicht: (2025)
von: Pavlíček, Martin, et al.
Veröffentlicht: (2025)
A multilingual hallucination benchmark: MultiWikiQHalluA
von: Thoresen, Freja, et al.
Veröffentlicht: (2026)
von: Thoresen, Freja, et al.
Veröffentlicht: (2026)
Suvach -- Generated Hindi QA benchmark
von: Narayanan, Vaishak, et al.
Veröffentlicht: (2024)
von: Narayanan, Vaishak, et al.
Veröffentlicht: (2024)
Evaluation and Facilitation of Online Discussions in the LLM Era: A Survey
von: Korre, Katerina, et al.
Veröffentlicht: (2025)
von: Korre, Katerina, et al.
Veröffentlicht: (2025)
LongTail-Swap: benchmarking language models' abilities on rare words
von: Algayres, Robin, et al.
Veröffentlicht: (2025)
von: Algayres, Robin, et al.
Veröffentlicht: (2025)
ConspEmoLLM-v2: A robust and stable model to detect sentiment-transformed conspiracy theories
von: Liu, Zhiwei, et al.
Veröffentlicht: (2025)
von: Liu, Zhiwei, et al.
Veröffentlicht: (2025)
LLM-MC-Affect: LLM-Based Monte Carlo Modeling of Affective Trajectories and Latent Ambiguity for Interpersonal Dynamic Insight
von: Lin, Yu-Zheng, et al.
Veröffentlicht: (2026)
von: Lin, Yu-Zheng, et al.
Veröffentlicht: (2026)
From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
von: Stephan, Andreas, et al.
Veröffentlicht: (2024)
von: Stephan, Andreas, et al.
Veröffentlicht: (2024)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
von: Phute, Mansi, et al.
Veröffentlicht: (2023)
von: Phute, Mansi, et al.
Veröffentlicht: (2023)
MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models
von: Yoshitake, Michiko, et al.
Veröffentlicht: (2026)
von: Yoshitake, Michiko, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
von: Siska, Charlotte, et al.
Veröffentlicht: (2025) -
From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
von: Roemmele, Melissa, et al.
Veröffentlicht: (2024) -
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
von: Liu, Shuyu, et al.
Veröffentlicht: (2025) -
Efficient Uncertainty Estimation for LLM-based Entity Linking in Tabular Data
von: Bono, Carlo, et al.
Veröffentlicht: (2025) -
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
von: Wen, Zichen, et al.
Veröffentlicht: (2024)