Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories
Fuente:
arXiv
Guardado en:
| Autores principales: | Cuesta-Ramirez, Jhouben, Beaussant, Samuel, Mounsif, Mehdi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scaling Algorithm Distillation for Continuous Control with Mamba
por: Beaussant, Samuel, et al.
Publicado: (2025)
por: Beaussant, Samuel, et al.
Publicado: (2025)
StyleBench: Evaluating thinking styles in Large Language Models
por: Guo, Junyu, et al.
Publicado: (2025)
por: Guo, Junyu, et al.
Publicado: (2025)
Can machines think efficiently?
por: Winchell, Adam
Publicado: (2025)
por: Winchell, Adam
Publicado: (2025)
Can Slow-thinking LLMs Reason Over Time? Empirical Studies in Time Series Forecasting
por: Cheng, Mingyue, et al.
Publicado: (2025)
por: Cheng, Mingyue, et al.
Publicado: (2025)
The boosted HP filter is more general than you might think
por: Mei, Ziwei, et al.
Publicado: (2022)
por: Mei, Ziwei, et al.
Publicado: (2022)
Causal thinking for decision making on Electronic Health Records: why and how
por: Doutreligne, Matthieu, et al.
Publicado: (2023)
por: Doutreligne, Matthieu, et al.
Publicado: (2023)
Replacing thinking with tool usage enables reasoning in small language models
por: Rainone, Corrado, et al.
Publicado: (2025)
por: Rainone, Corrado, et al.
Publicado: (2025)
Re-thinking Human Activity Recognition with Hierarchy-aware Label Relationship Modeling
por: Zuo, Jingwei, et al.
Publicado: (2024)
por: Zuo, Jingwei, et al.
Publicado: (2024)
Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured Spaces
por: Moss, Henry B., et al.
Publicado: (2025)
por: Moss, Henry B., et al.
Publicado: (2025)
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
por: Dutta, Subhabrata, et al.
Publicado: (2024)
por: Dutta, Subhabrata, et al.
Publicado: (2024)
Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pretraining
por: Li, Jeffrey, et al.
Publicado: (2026)
por: Li, Jeffrey, et al.
Publicado: (2026)
Real-valued continued fraction of straight lines
por: S, Vijay Prakash
Publicado: (2024)
por: S, Vijay Prakash
Publicado: (2024)
Making medical vision-language models think causally across modalities with retrieval-augmented cross-modal reasoning
por: Yang, Weiqin, et al.
Publicado: (2026)
por: Yang, Weiqin, et al.
Publicado: (2026)
ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
por: Wan, Ziyu, et al.
Publicado: (2025)
por: Wan, Ziyu, et al.
Publicado: (2025)
The 2020 US Decennial Census is more private than you (might) think
por: Su, Buxin, et al.
Publicado: (2024)
por: Su, Buxin, et al.
Publicado: (2024)
Is poisoning a real threat to LLM alignment? Maybe more so than you think
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2024)
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2024)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
por: Gor, Maharshi, et al.
Publicado: (2024)
por: Gor, Maharshi, et al.
Publicado: (2024)
Is Complex Training Necessary for Long-Tailed OOD Detection? A Re-think from Feature Geometry
por: Peng, Ningkang, et al.
Publicado: (2026)
por: Peng, Ningkang, et al.
Publicado: (2026)
Fitting magnetization data using continued fraction of straight lines
por: S, Vijay Prakash
Publicado: (2025)
por: S, Vijay Prakash
Publicado: (2025)
DINOv2 Rocks Geological Image Analysis: Classification, Segmentation, and Interpretability
por: Brondolo, Florent, et al.
Publicado: (2024)
por: Brondolo, Florent, et al.
Publicado: (2024)
Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models
por: Nezhurina, Marianna, et al.
Publicado: (2024)
por: Nezhurina, Marianna, et al.
Publicado: (2024)
Large Language Models as Analogical Reasoners
por: Yasunaga, Michihiro, et al.
Publicado: (2023)
por: Yasunaga, Michihiro, et al.
Publicado: (2023)
Can Large Reasoning Models Self-Train?
por: Shafayat, Sheikh, et al.
Publicado: (2025)
por: Shafayat, Sheikh, et al.
Publicado: (2025)
Decoding the Critique Mechanism in Large Reasoning Models
por: Phan, Hoang, et al.
Publicado: (2026)
por: Phan, Hoang, et al.
Publicado: (2026)
The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models
por: Anderson, Samuel Cyrenius
Publicado: (2026)
por: Anderson, Samuel Cyrenius
Publicado: (2026)
Are Large Reasoning Models Interruptible?
por: Wu, Tsung-Han, et al.
Publicado: (2025)
por: Wu, Tsung-Han, et al.
Publicado: (2025)
ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
por: Kwon, Yongchan, et al.
Publicado: (2025)
por: Kwon, Yongchan, et al.
Publicado: (2025)
The Impact of Quantization on Large Reasoning Model Reinforcement Learning
por: Kumar, Medha, et al.
Publicado: (2025)
por: Kumar, Medha, et al.
Publicado: (2025)
Teaching Large Language Models to Reason with Reinforcement Learning
por: Havrilla, Alex, et al.
Publicado: (2024)
por: Havrilla, Alex, et al.
Publicado: (2024)
ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models
por: Lucas, Ryan, et al.
Publicado: (2026)
por: Lucas, Ryan, et al.
Publicado: (2026)
Towards Large Reasoning Models for Agriculture
por: Zaremehrjerdi, Hossein, et al.
Publicado: (2025)
por: Zaremehrjerdi, Hossein, et al.
Publicado: (2025)
Confidence in the Reasoning of Large Language Models
por: Pawitan, Yudi, et al.
Publicado: (2024)
por: Pawitan, Yudi, et al.
Publicado: (2024)
Diversity-Aware Policy Optimization for Large Language Model Reasoning
por: Yao, Jian, et al.
Publicado: (2025)
por: Yao, Jian, et al.
Publicado: (2025)
Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models
por: Zhou, Zhanke, et al.
Publicado: (2025)
por: Zhou, Zhanke, et al.
Publicado: (2025)
TEMPO: Scaling Test-time Training for Large Reasoning Models
por: Zhang, Qingyang, et al.
Publicado: (2026)
por: Zhang, Qingyang, et al.
Publicado: (2026)
A Multi-task Large Reasoning Model for Molecular Science
por: Liu, Pengfei, et al.
Publicado: (2026)
por: Liu, Pengfei, et al.
Publicado: (2026)
Emergent Manifold Separability during Reasoning in Large Language Models
por: Chun, Chanwoo, et al.
Publicado: (2026)
por: Chun, Chanwoo, et al.
Publicado: (2026)
Re-thinking Adult Education Research. Beyond the Pandemic
Publicado: (2023)
Publicado: (2023)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
por: Huang, Jerry, et al.
Publicado: (2024)
por: Huang, Jerry, et al.
Publicado: (2024)
AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
por: Liang, Jiacheng, et al.
Publicado: (2025)
por: Liang, Jiacheng, et al.
Publicado: (2025)
Ejemplares similares
-
Scaling Algorithm Distillation for Continuous Control with Mamba
por: Beaussant, Samuel, et al.
Publicado: (2025) -
StyleBench: Evaluating thinking styles in Large Language Models
por: Guo, Junyu, et al.
Publicado: (2025) -
Can machines think efficiently?
por: Winchell, Adam
Publicado: (2025) -
Can Slow-thinking LLMs Reason Over Time? Empirical Studies in Time Series Forecasting
por: Cheng, Mingyue, et al.
Publicado: (2025) -
The boosted HP filter is more general than you might think
por: Mei, Ziwei, et al.
Publicado: (2022)