A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
Fuente:
arXiv
Saved in:
| Main Authors: | Hochlehnert, Andreas, Bhatnagar, Hardik, Udandarao, Vishaal, Albanie, Samuel, Prabhu, Ameya, Bethge, Matthias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
by: Ghosh, Adhiraj, et al.
Published: (2024)
by: Ghosh, Adhiraj, et al.
Published: (2024)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
How to Merge Your Multimodal Models Over Time?
by: Dziadzio, Sebastian, et al.
Published: (2024)
by: Dziadzio, Sebastian, et al.
Published: (2024)
Mapping Post-Training Forgetting in Language Models at Scale
by: Harmon, Jackson, et al.
Published: (2025)
by: Harmon, Jackson, et al.
Published: (2025)
CiteME: Can Language Models Accurately Cite Scientific Claims?
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
Solving Spatial Supersensing Without Spatial Supersensing
by: Udandarao, Vishaal, et al.
Published: (2025)
by: Udandarao, Vishaal, et al.
Published: (2025)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
A Practitioner's Guide to Continual Multimodal Pretraining
by: Roth, Karsten, et al.
Published: (2024)
by: Roth, Karsten, et al.
Published: (2024)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
by: Sinha, Shiven, et al.
Published: (2024)
by: Sinha, Shiven, et al.
Published: (2024)
Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs
by: Schuhmann, Christoph, et al.
Published: (2025)
by: Schuhmann, Christoph, et al.
Published: (2025)
LLM generation novelty through the lens of semantic similarity
by: Davydov, Philipp, et al.
Published: (2025)
by: Davydov, Philipp, et al.
Published: (2025)
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
by: Udandarao, Vishaal, et al.
Published: (2025)
by: Udandarao, Vishaal, et al.
Published: (2025)
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
by: Godbole, Ameya, et al.
Published: (2025)
by: Godbole, Ameya, et al.
Published: (2025)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026)
by: Rank, Ben, et al.
Published: (2026)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
by: Skorobogat, Ronald, et al.
Published: (2026)
by: Skorobogat, Ronald, et al.
Published: (2026)
Scaling Open-Ended Reasoning to Predict the Future
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Great Models Think Alike and this Undermines AI Oversight
by: Goel, Shashwat, et al.
Published: (2025)
by: Goel, Shashwat, et al.
Published: (2025)
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation
by: Sinha, Shiven, et al.
Published: (2025)
by: Sinha, Shiven, et al.
Published: (2025)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Are We Done with Object-Centric Learning?
by: Rubinstein, Alexander, et al.
Published: (2025)
by: Rubinstein, Alexander, et al.
Published: (2025)
Active Data Curation Effectively Distills Large-Scale Multimodal Models
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
Concept-Aware Batch Sampling Improves Language-Image Pretraining
by: Ghosh, Adhiraj, et al.
Published: (2025)
by: Ghosh, Adhiraj, et al.
Published: (2025)
Data-Centric Lessons To Improve Speech-Language Pretraining
by: Udandarao, Vishaal, et al.
Published: (2025)
by: Udandarao, Vishaal, et al.
Published: (2025)
PEDAL: Enhancing Greedy Decoding with Large Language Models using Diverse Exemplars
by: Prabhu, Sumanth
Published: (2024)
by: Prabhu, Sumanth
Published: (2024)
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
by: Wiedemer, Thaddäus, et al.
Published: (2025)
by: Wiedemer, Thaddäus, et al.
Published: (2025)
Progressive-Hint Prompting Improves Reasoning in Large Language Models
by: Zheng, Chuanyang, et al.
Published: (2023)
by: Zheng, Chuanyang, et al.
Published: (2023)
SoftLMs: Efficient Adaptive Low-Rank Approximation of Language Models using Soft-Thresholding Mechanism
by: Bhatnagar, Priyansh, et al.
Published: (2024)
by: Bhatnagar, Priyansh, et al.
Published: (2024)
A Sober Look at the Robustness of CLIPs to Spurious Features
by: Wang, Qizhou, et al.
Published: (2024)
by: Wang, Qizhou, et al.
Published: (2024)
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
by: Thede, Lukas, et al.
Published: (2025)
by: Thede, Lukas, et al.
Published: (2025)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
by: Roy, Tiasa Singha, et al.
Published: (2025)
by: Roy, Tiasa Singha, et al.
Published: (2025)
Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error Minimization
by: Shin, Sungbin, et al.
Published: (2024)
by: Shin, Sungbin, et al.
Published: (2024)
Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical Guidelines
by: Li, Yuchen, et al.
Published: (2024)
by: Li, Yuchen, et al.
Published: (2024)
Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
Mapping Faithful Reasoning in Language Models
by: Li, Jiazheng, et al.
Published: (2025)
by: Li, Jiazheng, et al.
Published: (2025)
SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning
by: Hu, Chenzhi, et al.
Published: (2026)
by: Hu, Chenzhi, et al.
Published: (2026)
Manifold-based Sampling for In-Context Hallucination Detection in Large Language Models
by: Vamshi, Bodla Krishna, et al.
Published: (2026)
by: Vamshi, Bodla Krishna, et al.
Published: (2026)
A Closer Look into Mixture-of-Experts in Large Language Models
by: Lo, Ka Man, et al.
Published: (2024)
by: Lo, Ka Man, et al.
Published: (2024)
Data Contamination Report from the 2024 CONDA Shared Task
by: Sainz, Oscar, et al.
Published: (2024)
by: Sainz, Oscar, et al.
Published: (2024)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement
by: Yuan, Hui, et al.
Published: (2024)
by: Yuan, Hui, et al.
Published: (2024)
Similar Items
-
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
by: Ghosh, Adhiraj, et al.
Published: (2024) -
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024) -
How to Merge Your Multimodal Models Over Time?
by: Dziadzio, Sebastian, et al.
Published: (2024) -
Mapping Post-Training Forgetting in Language Models at Scale
by: Harmon, Jackson, et al.
Published: (2025) -
CiteME: Can Language Models Accurately Cite Scientific Claims?
by: Press, Ori, et al.
Published: (2024)