Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mousavi, Seyed Mahed, Cecchinato, Edoardo, Hornikova, Lucia, Riccardi, Giuseppe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2026)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2026)
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2025)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2025)
DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
[De|Re]constructing VLMs' Reasoning in Counting
von: Alghisi, Simone, et al.
Veröffentlicht: (2025)
von: Alghisi, Simone, et al.
Veröffentlicht: (2025)
Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for Dialogue
von: Alghisi, Simone, et al.
Veröffentlicht: (2024)
von: Alghisi, Simone, et al.
Veröffentlicht: (2024)
Are LLMs Robust for Spoken Dialogues?
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
CIVET: Systematic Evaluation of Understanding in VLMs
von: Rizzoli, Massimo, et al.
Veröffentlicht: (2025)
von: Rizzoli, Massimo, et al.
Veröffentlicht: (2025)
Getting to the Point: Pointing Improves LVLMs at Counting
von: Alghisi, Simone, et al.
Veröffentlicht: (2026)
von: Alghisi, Simone, et al.
Veröffentlicht: (2026)
V-DyKnow: A Dynamic Benchmark for Time-Sensitive Knowledge in Vision Language Models
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2026)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2026)
MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs
von: Roccabruna, Gabriel, et al.
Veröffentlicht: (2026)
von: Roccabruna, Gabriel, et al.
Veröffentlicht: (2026)
What Are They Doing? Joint Audio-Speech Co-Reasoning
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024)
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024)
The Reasoning Error About Reasoning: Why Different Types of Reasoning Require Different Representational Structures
von: Wu, Yiling
Veröffentlicht: (2026)
von: Wu, Yiling
Veröffentlicht: (2026)
What Are They Talking About? A Benchmark of Knowledge-Grounded Discussion Summarization
von: Zhou, Weixiao, et al.
Veröffentlicht: (2025)
von: Zhou, Weixiao, et al.
Veröffentlicht: (2025)
Virtual Garbage Collector (VGC): A Zone-Based Garbage Collection Architecture for Python's Parallel Runtime
von: M, Abdulla
Veröffentlicht: (2025)
von: M, Abdulla
Veröffentlicht: (2025)
What Do Speech Foundation Models Not Learn About Speech?
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2024)
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2024)
Thinking Out Loud: Do Reasoning Models Know When They're Right?
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2025)
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2025)
What Are They Filtering Out? An Experimental Benchmark of Filtering Strategies for Harm Reduction in Pretraining Datasets
von: Stranisci, Marco Antonio, et al.
Veröffentlicht: (2025)
von: Stranisci, Marco Antonio, et al.
Veröffentlicht: (2025)
What Do AI Agents Talk About? Discourse and Architectural Constraints in the First AI-Only Social Network
von: Dube, Taksch, et al.
Veröffentlicht: (2026)
von: Dube, Taksch, et al.
Veröffentlicht: (2026)
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
von: Kang, Deokhyung, et al.
Veröffentlicht: (2025)
von: Kang, Deokhyung, et al.
Veröffentlicht: (2025)
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
von: Wu, Mingqi, et al.
Veröffentlicht: (2025)
von: Wu, Mingqi, et al.
Veröffentlicht: (2025)
Classifying Unreliable Narrators with Large Language Models
von: Brei, Anneliese, et al.
Veröffentlicht: (2025)
von: Brei, Anneliese, et al.
Veröffentlicht: (2025)
What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
von: Noels, Sander, et al.
Veröffentlicht: (2025)
von: Noels, Sander, et al.
Veröffentlicht: (2025)
What Do Self-Supervised Speech Models Know About Words?
von: Pasad, Ankita, et al.
Veröffentlicht: (2023)
von: Pasad, Ankita, et al.
Veröffentlicht: (2023)
Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning
von: Sok, Jaewon, et al.
Veröffentlicht: (2026)
von: Sok, Jaewon, et al.
Veröffentlicht: (2026)
Will LLMs Replace the Encoder-Only Models in Temporal Relation Classification?
von: Roccabruna, Gabriel, et al.
Veröffentlicht: (2024)
von: Roccabruna, Gabriel, et al.
Veröffentlicht: (2024)
HearSay Benchmark: Do Audio LLMs Leak What They Hear?
von: Wang, Jin, et al.
Veröffentlicht: (2026)
von: Wang, Jin, et al.
Veröffentlicht: (2026)
PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian
von: Mozafari, Jamshid, et al.
Veröffentlicht: (2026)
von: Mozafari, Jamshid, et al.
Veröffentlicht: (2026)
BeDiscovER: The Benchmark of Discourse Understanding in the Era of Reasoning Language Models
von: Li, Chuyuan, et al.
Veröffentlicht: (2025)
von: Li, Chuyuan, et al.
Veröffentlicht: (2025)
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
von: Chizhov, Pavel, et al.
Veröffentlicht: (2025)
von: Chizhov, Pavel, et al.
Veröffentlicht: (2025)
A Primer in Post-Training Reasoning Data: What We Know About How It Works
von: Li, Yaoming, et al.
Veröffentlicht: (2026)
von: Li, Yaoming, et al.
Veröffentlicht: (2026)
Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning
von: Damirchi, Hamed, et al.
Veröffentlicht: (2026)
von: Damirchi, Hamed, et al.
Veröffentlicht: (2026)
Do Reasoning LLMs Refuse What They Infer in Long Contexts?
von: Fu, Yu, et al.
Veröffentlicht: (2026)
von: Fu, Yu, et al.
Veröffentlicht: (2026)
Reasoning Models Will Sometimes Lie About Their Reasoning
von: Walden, William, et al.
Veröffentlicht: (2026)
von: Walden, William, et al.
Veröffentlicht: (2026)
What Do LLMs Know About Alzheimer's Disease? Multi-loss Fine-Tuning and Probing for AD Detection
von: Jiang, Lei, et al.
Veröffentlicht: (2026)
von: Jiang, Lei, et al.
Veröffentlicht: (2026)
Towards Efficient Patient Recruitment for Clinical Trials: Application of a Prompt-Based Learning Model
von: Rahmanian, Mojdeh, et al.
Veröffentlicht: (2024)
von: Rahmanian, Mojdeh, et al.
Veröffentlicht: (2024)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
von: Mundada, Gagan, et al.
Veröffentlicht: (2025)
von: Mundada, Gagan, et al.
Veröffentlicht: (2025)
Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented Generation
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
It Is Not About What You Say, It Is About How You Say It: A Surprisingly Simple Approach for Improving Reading Comprehension
von: Shaier, Sagi, et al.
Veröffentlicht: (2024)
von: Shaier, Sagi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2026) -
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2025) -
DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024) -
[De|Re]constructing VLMs' Reasoning in Counting
von: Alghisi, Simone, et al.
Veröffentlicht: (2025) -
Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for Dialogue
von: Alghisi, Simone, et al.
Veröffentlicht: (2024)