ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Potamitis, Nearchos, Ramani, Vansh, Arora, Har Ashish, Kuchhal, Dhairya, Klein, Lars, Arora, Akhil |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Are Retrials All You Need? Enhancing Large Language Model Reasoning Without Verbalized Feedback
por: Potamitis, Nearchos, et al.
Publicado: (2025)
por: Potamitis, Nearchos, et al.
Publicado: (2025)
Fleet of Agents: Coordinated Problem Solving with Large Language Models
por: Klein, Lars, et al.
Publicado: (2024)
por: Klein, Lars, et al.
Publicado: (2024)
Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows
por: Mohammadi, Bardia, et al.
Publicado: (2026)
por: Mohammadi, Bardia, et al.
Publicado: (2026)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
por: Cheng, Yun, et al.
Publicado: (2026)
por: Cheng, Yun, et al.
Publicado: (2026)
MoBayes: A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support
por: Kesmen, Yusuf, et al.
Publicado: (2026)
por: Kesmen, Yusuf, et al.
Publicado: (2026)
RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
por: Abhyankar, Nikhil, et al.
Publicado: (2025)
por: Abhyankar, Nikhil, et al.
Publicado: (2025)
DIVERGE: Diversity-Enhanced RAG for Open-Ended Information Seeking
por: Hu, Tianyi, et al.
Publicado: (2026)
por: Hu, Tianyi, et al.
Publicado: (2026)
Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities
por: Zhao, Haoyu, et al.
Publicado: (2025)
por: Zhao, Haoyu, et al.
Publicado: (2025)
CosmicFish-HRM: Adaptive Reasoning via Hierarchical Recurrent Mechanisms in Compact Language Models
por: Lakkapragada, Venkat Akhil
Publicado: (2026)
por: Lakkapragada, Venkat Akhil
Publicado: (2026)
TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks
por: Kapoor, Vansh, et al.
Publicado: (2026)
por: Kapoor, Vansh, et al.
Publicado: (2026)
Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools
por: Mohammadi, Bardia, et al.
Publicado: (2026)
por: Mohammadi, Bardia, et al.
Publicado: (2026)
On the Reasoning Abilities of Masked Diffusion Language Models
por: Svete, Anej, et al.
Publicado: (2025)
por: Svete, Anej, et al.
Publicado: (2025)
Evaluating the Effectiveness of Data Augmentation for Emotion Classification in Low-Resource Settings
por: Arora, Aashish, et al.
Publicado: (2024)
por: Arora, Aashish, et al.
Publicado: (2024)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
por: Oh, Jungwoo, et al.
Publicado: (2026)
por: Oh, Jungwoo, et al.
Publicado: (2026)
Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
por: Kawakami, Wataru, et al.
Publicado: (2025)
por: Kawakami, Wataru, et al.
Publicado: (2025)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
por: Wei, Anjiang, et al.
Publicado: (2025)
por: Wei, Anjiang, et al.
Publicado: (2025)
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
por: Yang, Minglai, et al.
Publicado: (2025)
por: Yang, Minglai, et al.
Publicado: (2025)
Benchmarking ChatGPT on Algorithmic Reasoning
por: McLeish, Sean, et al.
Publicado: (2024)
por: McLeish, Sean, et al.
Publicado: (2024)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
por: Liu, Qihao, et al.
Publicado: (2025)
por: Liu, Qihao, et al.
Publicado: (2025)
ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
por: Yang, Ling, et al.
Publicado: (2025)
por: Yang, Ling, et al.
Publicado: (2025)
Weight-of-Thought Reasoning: Exploring Neural Network Weights for Enhanced LLM Reasoning
por: Punjwani, Saif, et al.
Publicado: (2025)
por: Punjwani, Saif, et al.
Publicado: (2025)
Explainable LLM Unlearning Through Reasoning
por: Liao, Junfeng, et al.
Publicado: (2026)
por: Liao, Junfeng, et al.
Publicado: (2026)
Token-Budget-Aware LLM Reasoning
por: Han, Tingxu, et al.
Publicado: (2024)
por: Han, Tingxu, et al.
Publicado: (2024)
SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
por: Zhao, Zehua, et al.
Publicado: (2025)
por: Zhao, Zehua, et al.
Publicado: (2025)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
por: Yoa, Seungdong, et al.
Publicado: (2026)
por: Yoa, Seungdong, et al.
Publicado: (2026)
The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning
por: Shin, Kwan Soo
Publicado: (2026)
por: Shin, Kwan Soo
Publicado: (2026)
Why is Your Language Model a Poor Implicit Reward Model?
por: Razin, Noam, et al.
Publicado: (2025)
por: Razin, Noam, et al.
Publicado: (2025)
Better LLM Reasoning via Dual-Play
por: Zhang, Zhengxin, et al.
Publicado: (2025)
por: Zhang, Zhengxin, et al.
Publicado: (2025)
The Impact of Language Mixing on Bilingual LLM Reasoning
por: Li, Yihao, et al.
Publicado: (2025)
por: Li, Yihao, et al.
Publicado: (2025)
Advancing LLM Reasoning Generalists with Preference Trees
por: Yuan, Lifan, et al.
Publicado: (2024)
por: Yuan, Lifan, et al.
Publicado: (2024)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
por: Lin, Zicheng, et al.
Publicado: (2024)
por: Lin, Zicheng, et al.
Publicado: (2024)
Directional Attractors in LLM Reasoning: How Similarity Retrieval Steers Iterative Summarization Based Reasoning
por: Tekin, Cagatay, et al.
Publicado: (2025)
por: Tekin, Cagatay, et al.
Publicado: (2025)
Training Language Models to Reason Efficiently
por: Arora, Daman, et al.
Publicado: (2025)
por: Arora, Daman, et al.
Publicado: (2025)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
por: Lin, Bill Yuchen, et al.
Publicado: (2025)
por: Lin, Bill Yuchen, et al.
Publicado: (2025)
On the Optimizer Dependence of Neural Scaling Laws
por: Ramani, Vansh, et al.
Publicado: (2026)
por: Ramani, Vansh, et al.
Publicado: (2026)
From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
por: Jiang, Xitai, et al.
Publicado: (2026)
por: Jiang, Xitai, et al.
Publicado: (2026)
Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia
por: Feith, Tomás, et al.
Publicado: (2024)
por: Feith, Tomás, et al.
Publicado: (2024)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
por: Didolkar, Aniket, et al.
Publicado: (2025)
por: Didolkar, Aniket, et al.
Publicado: (2025)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
por: Dong, Harry, et al.
Publicado: (2025)
por: Dong, Harry, et al.
Publicado: (2025)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
por: Macar, Uzay, et al.
Publicado: (2025)
por: Macar, Uzay, et al.
Publicado: (2025)
Ejemplares similares
-
Are Retrials All You Need? Enhancing Large Language Model Reasoning Without Verbalized Feedback
por: Potamitis, Nearchos, et al.
Publicado: (2025) -
Fleet of Agents: Coordinated Problem Solving with Large Language Models
por: Klein, Lars, et al.
Publicado: (2024) -
Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows
por: Mohammadi, Bardia, et al.
Publicado: (2026) -
Contextual Drag: How Errors in the Context Affect LLM Reasoning
por: Cheng, Yun, et al.
Publicado: (2026) -
MoBayes: A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support
por: Kesmen, Yusuf, et al.
Publicado: (2026)