Brittlebench: Quantifying LLM robustness via prompt sensitivity
Fuente:
arXiv
Guardado en:
| Autores principales: | Romanou, Angelika, Ibrahim, Mark, Ross, Candace, Shaib, Chantal, Oktar, Kerem, Bell, Samuel J., Ovalle, Anaelia, Dodge, Jesse, Bosselut, Antoine, Sinha, Koustuv, Williams, Adina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
por: Ovalle, Anaelia, et al.
Publicado: (2025)
por: Ovalle, Anaelia, et al.
Publicado: (2025)
PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
por: Chen, Zeming, et al.
Publicado: (2025)
por: Chen, Zeming, et al.
Publicado: (2025)
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
por: Bhagwatkar, Rishika, et al.
Publicado: (2025)
por: Bhagwatkar, Rishika, et al.
Publicado: (2025)
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
por: Ross, Candace, et al.
Publicado: (2025)
por: Ross, Candace, et al.
Publicado: (2025)
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
por: Ross, Candace, et al.
Publicado: (2024)
por: Ross, Candace, et al.
Publicado: (2024)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
por: Hall, Melissa, et al.
Publicado: (2024)
por: Hall, Melissa, et al.
Publicado: (2024)
Eval Factsheets: A Structured Framework for Documenting AI Evaluations
por: Bordes, Florian, et al.
Publicado: (2025)
por: Bordes, Florian, et al.
Publicado: (2025)
Changing Answer Order Can Decrease MMLU Accuracy
por: Gupta, Vipul, et al.
Publicado: (2024)
por: Gupta, Vipul, et al.
Publicado: (2024)
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
por: Kambhatla, Gauri, et al.
Publicado: (2025)
por: Kambhatla, Gauri, et al.
Publicado: (2025)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
por: Krojer, Benno, et al.
Publicado: (2025)
por: Krojer, Benno, et al.
Publicado: (2025)
The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models
por: Ovalle, Anaelia, et al.
Publicado: (2024)
por: Ovalle, Anaelia, et al.
Publicado: (2024)
Multi-Modal Language Models as Text-to-Image Model Evaluators
por: Chen, Jiahui, et al.
Publicado: (2025)
por: Chen, Jiahui, et al.
Publicado: (2025)
A computing machinery using a continuous memory tape
por: Oktar, Yigit
Publicado: (2023)
por: Oktar, Yigit
Publicado: (2023)
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
por: Hall, Melissa, et al.
Publicado: (2023)
por: Hall, Melissa, et al.
Publicado: (2023)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
por: Gupta, Vipul, et al.
Publicado: (2024)
por: Gupta, Vipul, et al.
Publicado: (2024)
Who Taught You That? Tracing Teachers in Model Distillation
por: Wadhwa, Somin, et al.
Publicado: (2025)
por: Wadhwa, Somin, et al.
Publicado: (2025)
Do different prompting methods yield a common task representation in language models?
por: Davidson, Guy, et al.
Publicado: (2025)
por: Davidson, Guy, et al.
Publicado: (2025)
LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling
por: Mondorf, Philipp, et al.
Publicado: (2026)
por: Mondorf, Philipp, et al.
Publicado: (2026)
Measuring AI "Slop" in Text
por: Shaib, Chantal, et al.
Publicado: (2025)
por: Shaib, Chantal, et al.
Publicado: (2025)
Detection and Measurement of Syntactic Templates in Generated Text
por: Shaib, Chantal, et al.
Publicado: (2024)
por: Shaib, Chantal, et al.
Publicado: (2024)
Efficient Tool Use with Chain-of-Abstraction Reasoning
por: Gao, Silin, et al.
Publicado: (2024)
por: Gao, Silin, et al.
Publicado: (2024)
Are Large Language Models Sensitive to the Motives Behind Communication?
por: Wu, Addison J., et al.
Publicado: (2025)
por: Wu, Addison J., et al.
Publicado: (2025)
RLMEval: Evaluating Research-Level Neural Theorem Proving
por: Poiroux, Auguste, et al.
Publicado: (2025)
por: Poiroux, Auguste, et al.
Publicado: (2025)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
por: Kim, Kyuhee, et al.
Publicado: (2026)
por: Kim, Kyuhee, et al.
Publicado: (2026)
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
por: Bayazit, Deniz, et al.
Publicado: (2025)
por: Bayazit, Deniz, et al.
Publicado: (2025)
Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
por: Shaib, Chantal, et al.
Publicado: (2025)
por: Shaib, Chantal, et al.
Publicado: (2025)
How Much Annotation is Needed to Compare Summarization Models?
por: Shaib, Chantal, et al.
Publicado: (2024)
por: Shaib, Chantal, et al.
Publicado: (2024)
QA-prompting: Improving Summarization with Large Language Models using Question-Answering
por: Sinha, Neelabh
Publicado: (2025)
por: Sinha, Neelabh
Publicado: (2025)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
por: Fan, Dongyang, et al.
Publicado: (2025)
por: Fan, Dongyang, et al.
Publicado: (2025)
SCRuB: Social Concept Reasoning under Rubric-Based Evaluation
por: Watson-Daniels, Jamelle, et al.
Publicado: (2026)
por: Watson-Daniels, Jamelle, et al.
Publicado: (2026)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
por: Foroutan, Negar, et al.
Publicado: (2025)
por: Foroutan, Negar, et al.
Publicado: (2025)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
por: Mañas, Oscar, et al.
Publicado: (2024)
por: Mañas, Oscar, et al.
Publicado: (2024)
Emergent misalignment as prompt sensitivity: A research note
por: Wyse, Tim, et al.
Publicado: (2025)
por: Wyse, Tim, et al.
Publicado: (2025)
Identifying, Evaluating, and Mitigating Risks of AI Thought Partnerships
por: Oktar, Kerem, et al.
Publicado: (2025)
por: Oktar, Kerem, et al.
Publicado: (2025)
JOBSKAPE: A Framework for Generating Synthetic Job Postings to Enhance Skill Matching
por: Magron, Antoine, et al.
Publicado: (2024)
por: Magron, Antoine, et al.
Publicado: (2024)
The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation
por: Mundt, Martin, et al.
Publicado: (2025)
por: Mundt, Martin, et al.
Publicado: (2025)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
por: Gao, Silin, et al.
Publicado: (2025)
por: Gao, Silin, et al.
Publicado: (2025)
AI Meets Mathematics Education: A Case Study on Supporting an Instructor in a Large Mathematics Class with Context-Aware AI
por: Barghorn, Jérémy, et al.
Publicado: (2026)
por: Barghorn, Jérémy, et al.
Publicado: (2026)
Let Me Teach You: Pedagogical Foundations of Feedback for Language Models
por: Borges, Beatriz, et al.
Publicado: (2023)
por: Borges, Beatriz, et al.
Publicado: (2023)
Complex Reasoning over Logical Queries on Commonsense Knowledge Graphs
por: Fang, Tianqing, et al.
Publicado: (2024)
por: Fang, Tianqing, et al.
Publicado: (2024)
Ejemplares similares
-
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
por: Ovalle, Anaelia, et al.
Publicado: (2025) -
PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
por: Chen, Zeming, et al.
Publicado: (2025) -
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
por: Bhagwatkar, Rishika, et al.
Publicado: (2025) -
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
por: Ross, Candace, et al.
Publicado: (2025) -
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
por: Ross, Candace, et al.
Publicado: (2024)