Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | Ailem, Melissa, Marazopoulou, Katerina, Siska, Charlotte, Bono, James |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
por: Siska, Charlotte, et al.
Publicado: (2025)
por: Siska, Charlotte, et al.
Publicado: (2025)
From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
por: Roemmele, Melissa, et al.
Publicado: (2024)
por: Roemmele, Melissa, et al.
Publicado: (2024)
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
por: Liu, Shuyu, et al.
Publicado: (2025)
por: Liu, Shuyu, et al.
Publicado: (2025)
Efficient Uncertainty Estimation for LLM-based Entity Linking in Tabular Data
por: Bono, Carlo, et al.
Publicado: (2025)
por: Bono, Carlo, et al.
Publicado: (2025)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
por: Wen, Zichen, et al.
Publicado: (2024)
por: Wen, Zichen, et al.
Publicado: (2024)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
por: Li, Kunning, et al.
Publicado: (2025)
por: Li, Kunning, et al.
Publicado: (2025)
CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance
por: Zhang, Feng, et al.
Publicado: (2025)
por: Zhang, Feng, et al.
Publicado: (2025)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
por: Montalan, Jann Railey, et al.
Publicado: (2025)
por: Montalan, Jann Railey, et al.
Publicado: (2025)
AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
por: Schmidgall, Samuel, et al.
Publicado: (2024)
por: Schmidgall, Samuel, et al.
Publicado: (2024)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
por: Jassim, Serwan, et al.
Publicado: (2023)
por: Jassim, Serwan, et al.
Publicado: (2023)
Dynamic benchmarking framework for LLM-based conversational data capture
por: Aluffi, Pietro Alessandro, et al.
Publicado: (2025)
por: Aluffi, Pietro Alessandro, et al.
Publicado: (2025)
LLMzSzŁ: a comprehensive LLM benchmark for Polish
por: Jassem, Krzysztof, et al.
Publicado: (2025)
por: Jassem, Krzysztof, et al.
Publicado: (2025)
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
por: Kankowski, Florian, et al.
Publicado: (2025)
por: Kankowski, Florian, et al.
Publicado: (2025)
OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities
por: Verma, Sahil, et al.
Publicado: (2025)
por: Verma, Sahil, et al.
Publicado: (2025)
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
por: Panagoulias, Dimitrios P., et al.
Publicado: (2024)
por: Panagoulias, Dimitrios P., et al.
Publicado: (2024)
Examining Identity Drift in Conversations of LLM Agents
por: Choi, Junhyuk, et al.
Publicado: (2024)
por: Choi, Junhyuk, et al.
Publicado: (2024)
Verbalizing LLMs' assumptions to explain and control sycophancy
por: Cheng, Myra, et al.
Publicado: (2026)
por: Cheng, Myra, et al.
Publicado: (2026)
Cross-lingual robustness of LLM-brain alignment and its computational roots
por: Yang, Ni, et al.
Publicado: (2026)
por: Yang, Ni, et al.
Publicado: (2026)
Monte Carlo Temperature: a robust sampling strategy for LLM's uncertainty quantification methods
por: Cecere, Nicola, et al.
Publicado: (2025)
por: Cecere, Nicola, et al.
Publicado: (2025)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
Examining the Behavior of LLM Architectures Within the Framework of Standardized National Exams in Brazil
por: Locatelli, Marcelo Sartori, et al.
Publicado: (2024)
por: Locatelli, Marcelo Sartori, et al.
Publicado: (2024)
Linguini: A benchmark for language-agnostic linguistic reasoning
por: Sánchez, Eduardo, et al.
Publicado: (2024)
por: Sánchez, Eduardo, et al.
Publicado: (2024)
An Expert-grounded benchmark of General Purpose LLMs in LCA
por: Donaldson, Artur, et al.
Publicado: (2025)
por: Donaldson, Artur, et al.
Publicado: (2025)
How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives
por: Ichmoukhamedov, Timour, et al.
Publicado: (2024)
por: Ichmoukhamedov, Timour, et al.
Publicado: (2024)
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples
por: Michail, Andrianos, et al.
Publicado: (2025)
por: Michail, Andrianos, et al.
Publicado: (2025)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
por: Almeida, Thales Sales, et al.
Publicado: (2025)
por: Almeida, Thales Sales, et al.
Publicado: (2025)
Certainty robustness: Evaluating LLM stability under self-challenging prompts
por: Saadat, Mohammadreza, et al.
Publicado: (2026)
por: Saadat, Mohammadreza, et al.
Publicado: (2026)
GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
por: Bandooni, Ashutosh, et al.
Publicado: (2025)
por: Bandooni, Ashutosh, et al.
Publicado: (2025)
NLP for The Greek Language: A Longer Survey
por: Papantoniou, Katerina, et al.
Publicado: (2024)
por: Papantoniou, Katerina, et al.
Publicado: (2024)
Proverbs or Pythian Oracles? Sentiments and Emotions in Greek Sayings
por: Korre, Katerina, et al.
Publicado: (2025)
por: Korre, Katerina, et al.
Publicado: (2025)
BIPOLAR: Polarization-based granular framework for LLM bias evaluation
por: Pavlíček, Martin, et al.
Publicado: (2025)
por: Pavlíček, Martin, et al.
Publicado: (2025)
A multilingual hallucination benchmark: MultiWikiQHalluA
por: Thoresen, Freja, et al.
Publicado: (2026)
por: Thoresen, Freja, et al.
Publicado: (2026)
Suvach -- Generated Hindi QA benchmark
por: Narayanan, Vaishak, et al.
Publicado: (2024)
por: Narayanan, Vaishak, et al.
Publicado: (2024)
Evaluation and Facilitation of Online Discussions in the LLM Era: A Survey
por: Korre, Katerina, et al.
Publicado: (2025)
por: Korre, Katerina, et al.
Publicado: (2025)
LongTail-Swap: benchmarking language models' abilities on rare words
por: Algayres, Robin, et al.
Publicado: (2025)
por: Algayres, Robin, et al.
Publicado: (2025)
ConspEmoLLM-v2: A robust and stable model to detect sentiment-transformed conspiracy theories
por: Liu, Zhiwei, et al.
Publicado: (2025)
por: Liu, Zhiwei, et al.
Publicado: (2025)
LLM-MC-Affect: LLM-Based Monte Carlo Modeling of Affective Trajectories and Latent Ambiguity for Interpersonal Dynamic Insight
por: Lin, Yu-Zheng, et al.
Publicado: (2026)
por: Lin, Yu-Zheng, et al.
Publicado: (2026)
From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
por: Stephan, Andreas, et al.
Publicado: (2024)
por: Stephan, Andreas, et al.
Publicado: (2024)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
por: Phute, Mansi, et al.
Publicado: (2023)
por: Phute, Mansi, et al.
Publicado: (2023)
MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models
por: Yoshitake, Michiko, et al.
Publicado: (2026)
por: Yoshitake, Michiko, et al.
Publicado: (2026)
Ejemplares similares
-
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
por: Siska, Charlotte, et al.
Publicado: (2025) -
From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
por: Roemmele, Melissa, et al.
Publicado: (2024) -
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
por: Liu, Shuyu, et al.
Publicado: (2025) -
Efficient Uncertainty Estimation for LLM-based Entity Linking in Tabular Data
por: Bono, Carlo, et al.
Publicado: (2025) -
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
por: Wen, Zichen, et al.
Publicado: (2024)