Saved in:
| Main Authors: | Romanou, Angelika, Ibrahim, Mark, Ross, Candace, Shaib, Chantal, Oktar, Kerem, Bell, Samuel J., Ovalle, Anaelia, Dodge, Jesse, Bosselut, Antoine, Sinha, Koustuv, Williams, Adina |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.13285 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
by: Ovalle, Anaelia, et al.
Published: (2025)
by: Ovalle, Anaelia, et al.
Published: (2025)
PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
by: Chen, Zeming, et al.
Published: (2025)
by: Chen, Zeming, et al.
Published: (2025)
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
by: Bhagwatkar, Rishika, et al.
Published: (2025)
by: Bhagwatkar, Rishika, et al.
Published: (2025)
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
by: Ross, Candace, et al.
Published: (2025)
by: Ross, Candace, et al.
Published: (2025)
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
by: Ross, Candace, et al.
Published: (2024)
by: Ross, Candace, et al.
Published: (2024)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
by: Hall, Melissa, et al.
Published: (2024)
by: Hall, Melissa, et al.
Published: (2024)
Eval Factsheets: A Structured Framework for Documenting AI Evaluations
by: Bordes, Florian, et al.
Published: (2025)
by: Bordes, Florian, et al.
Published: (2025)
Changing Answer Order Can Decrease MMLU Accuracy
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
by: Kambhatla, Gauri, et al.
Published: (2025)
by: Kambhatla, Gauri, et al.
Published: (2025)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
by: Krojer, Benno, et al.
Published: (2025)
by: Krojer, Benno, et al.
Published: (2025)
The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models
by: Ovalle, Anaelia, et al.
Published: (2024)
by: Ovalle, Anaelia, et al.
Published: (2024)
Multi-Modal Language Models as Text-to-Image Model Evaluators
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
by: Hall, Melissa, et al.
Published: (2023)
by: Hall, Melissa, et al.
Published: (2023)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
A computing machinery using a continuous memory tape
by: Oktar, Yigit
Published: (2023)
by: Oktar, Yigit
Published: (2023)
Who Taught You That? Tracing Teachers in Model Distillation
by: Wadhwa, Somin, et al.
Published: (2025)
by: Wadhwa, Somin, et al.
Published: (2025)
Measuring AI "Slop" in Text
by: Shaib, Chantal, et al.
Published: (2025)
by: Shaib, Chantal, et al.
Published: (2025)
Detection and Measurement of Syntactic Templates in Generated Text
by: Shaib, Chantal, et al.
Published: (2024)
by: Shaib, Chantal, et al.
Published: (2024)
LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling
by: Mondorf, Philipp, et al.
Published: (2026)
by: Mondorf, Philipp, et al.
Published: (2026)
Efficient Tool Use with Chain-of-Abstraction Reasoning
by: Gao, Silin, et al.
Published: (2024)
by: Gao, Silin, et al.
Published: (2024)
Do different prompting methods yield a common task representation in language models?
by: Davidson, Guy, et al.
Published: (2025)
by: Davidson, Guy, et al.
Published: (2025)
Are Large Language Models Sensitive to the Motives Behind Communication?
by: Wu, Addison J., et al.
Published: (2025)
by: Wu, Addison J., et al.
Published: (2025)
SCRuB: Social Concept Reasoning under Rubric-Based Evaluation
by: Watson-Daniels, Jamelle, et al.
Published: (2026)
by: Watson-Daniels, Jamelle, et al.
Published: (2026)
Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
by: Shaib, Chantal, et al.
Published: (2025)
by: Shaib, Chantal, et al.
Published: (2025)
How Much Annotation is Needed to Compare Summarization Models?
by: Shaib, Chantal, et al.
Published: (2024)
by: Shaib, Chantal, et al.
Published: (2024)
RLMEval: Evaluating Research-Level Neural Theorem Proving
by: Poiroux, Auguste, et al.
Published: (2025)
by: Poiroux, Auguste, et al.
Published: (2025)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
by: Kim, Kyuhee, et al.
Published: (2026)
by: Kim, Kyuhee, et al.
Published: (2026)
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
by: Bayazit, Deniz, et al.
Published: (2025)
by: Bayazit, Deniz, et al.
Published: (2025)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
by: Mañas, Oscar, et al.
Published: (2024)
by: Mañas, Oscar, et al.
Published: (2024)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation
by: Mundt, Martin, et al.
Published: (2025)
by: Mundt, Martin, et al.
Published: (2025)
Identifying, Evaluating, and Mitigating Risks of AI Thought Partnerships
by: Oktar, Kerem, et al.
Published: (2025)
by: Oktar, Kerem, et al.
Published: (2025)
QA-prompting: Improving Summarization with Large Language Models using Question-Answering
by: Sinha, Neelabh
Published: (2025)
by: Sinha, Neelabh
Published: (2025)
[Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
JOBSKAPE: A Framework for Generating Synthetic Job Postings to Enhance Skill Matching
by: Magron, Antoine, et al.
Published: (2024)
by: Magron, Antoine, et al.
Published: (2024)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
by: Gao, Silin, et al.
Published: (2025)
by: Gao, Silin, et al.
Published: (2025)
AI Meets Mathematics Education: A Case Study on Supporting an Instructor in a Large Mathematics Class with Context-Aware AI
by: Barghorn, Jérémy, et al.
Published: (2026)
by: Barghorn, Jérémy, et al.
Published: (2026)
Let Me Teach You: Pedagogical Foundations of Feedback for Language Models
by: Borges, Beatriz, et al.
Published: (2023)
by: Borges, Beatriz, et al.
Published: (2023)
Complex Reasoning over Logical Queries on Commonsense Knowledge Graphs
by: Fang, Tianqing, et al.
Published: (2024)
by: Fang, Tianqing, et al.
Published: (2024)
Similar Items
-
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
by: Ovalle, Anaelia, et al.
Published: (2025) -
PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
by: Chen, Zeming, et al.
Published: (2025) -
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
by: Bhagwatkar, Rishika, et al.
Published: (2025) -
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
by: Ross, Candace, et al.
Published: (2025) -
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
by: Ross, Candace, et al.
Published: (2024)