A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Qianjun, Ji, Wenkai, Ding, Yuyang, Li, Junsong, Chen, Shilian, Wang, Junyi, Zhou, Jie, Chen, Qin, Zhang, Min, Wu, Yulan, He, Liang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
von: Li, Junsong, et al.
Veröffentlicht: (2025)
von: Li, Junsong, et al.
Veröffentlicht: (2025)
Reinforced Interactive Continual Learning via Real-time Noisy Human Feedback
von: Yang, Yutao, et al.
Veröffentlicht: (2025)
von: Yang, Yutao, et al.
Veröffentlicht: (2025)
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
von: Li, Junsong, et al.
Veröffentlicht: (2025)
von: Li, Junsong, et al.
Veröffentlicht: (2025)
Black-box Model Merging for Language-Model-as-a-Service with Massive Model Repositories
von: Chen, Shilian, et al.
Veröffentlicht: (2025)
von: Chen, Shilian, et al.
Veröffentlicht: (2025)
Forget What's Sensitive, Remember What Matters: Token-Level Differential Privacy in Memory Sculpting for Continual Learning
von: Zhan, Bihao, et al.
Veröffentlicht: (2025)
von: Zhan, Bihao, et al.
Veröffentlicht: (2025)
Time Series Forecasting as Reasoning: A Slow-Thinking Approach with Reinforced LLMs
von: Zhou, Yitong, et al.
Veröffentlicht: (2025)
von: Zhou, Yitong, et al.
Veröffentlicht: (2025)
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
von: Paliotta, Daniele, et al.
Veröffentlicht: (2025)
von: Paliotta, Daniele, et al.
Veröffentlicht: (2025)
Emergent Slow Thinking in LLMs as Inverse Tree Freezing
von: Hu, Sihan, et al.
Veröffentlicht: (2025)
von: Hu, Sihan, et al.
Veröffentlicht: (2025)
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking
von: Cheng, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Cheng, Xiaoxue, et al.
Veröffentlicht: (2025)
PsychAgent: An Experience-Driven Lifelong Learning Agent for Self-Evolving Psychological Counselor
von: Yang, Yutao, et al.
Veröffentlicht: (2026)
von: Yang, Yutao, et al.
Veröffentlicht: (2026)
AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
von: Xiang, Kun, et al.
Veröffentlicht: (2024)
von: Xiang, Kun, et al.
Veröffentlicht: (2024)
From Sufficiency to Reflection: Reinforcement-Guided Thinking Quality in Retrieval-Augmented Reasoning for LLMs
von: He, Jie, et al.
Veröffentlicht: (2025)
von: He, Jie, et al.
Veröffentlicht: (2025)
Mathematical Language Models: A Survey
von: Liu, Wentao, et al.
Veröffentlicht: (2023)
von: Liu, Wentao, et al.
Veröffentlicht: (2023)
PsychEval: A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor
von: Pan, Qianjun, et al.
Veröffentlicht: (2026)
von: Pan, Qianjun, et al.
Veröffentlicht: (2026)
Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
von: Shen, Junhao, et al.
Veröffentlicht: (2025)
von: Shen, Junhao, et al.
Veröffentlicht: (2025)
Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking
von: Tian, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Tian, Xiaoyu, et al.
Veröffentlicht: (2025)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
von: Liu, Wentao, et al.
Veröffentlicht: (2024)
von: Liu, Wentao, et al.
Veröffentlicht: (2024)
Slow Thinking for Sequential Recommendation
von: Zhang, Junjie, et al.
Veröffentlicht: (2025)
von: Zhang, Junjie, et al.
Veröffentlicht: (2025)
From Prediction to Justification: Aligning Sentiment Reasoning with Human Rationale via Reinforcement Learning
von: Zhang, Shihao, et al.
Veröffentlicht: (2026)
von: Zhang, Shihao, et al.
Veröffentlicht: (2026)
AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
von: Yang, Yutao, et al.
Veröffentlicht: (2026)
von: Yang, Yutao, et al.
Veröffentlicht: (2026)
Identity, Crimes, and Law Enforcement in the Metaverse
von: Qin, Hua Xuan, et al.
Veröffentlicht: (2022)
von: Qin, Hua Xuan, et al.
Veröffentlicht: (2022)
Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality
von: Bi, Ziqian, et al.
Veröffentlicht: (2025)
von: Bi, Ziqian, et al.
Veröffentlicht: (2025)
Can Heterogeneous Language Models Be Fused?
von: Chen, Shilian, et al.
Veröffentlicht: (2026)
von: Chen, Shilian, et al.
Veröffentlicht: (2026)
Thinker: Learning to Think Fast and Slow
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning
von: Zhang, Xue, et al.
Veröffentlicht: (2025)
von: Zhang, Xue, et al.
Veröffentlicht: (2025)
Scaling Laws for Predicting Downstream Performance in LLMs
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
Unlocking a New Rust Programming Experience: Fast and Slow Thinking with LLMs to Conquer Undefined Behaviors
von: Jiang, Renshuang, et al.
Veröffentlicht: (2025)
von: Jiang, Renshuang, et al.
Veröffentlicht: (2025)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
Unsupervised Equivalent Contrastive Learning for Radio Signal Recognition
von: Zheng, Shilian, et al.
Veröffentlicht: (2026)
von: Zheng, Shilian, et al.
Veröffentlicht: (2026)
Relative Scaling Laws for LLMs
von: Held, William, et al.
Veröffentlicht: (2025)
von: Held, William, et al.
Veröffentlicht: (2025)
Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
von: Jia, Xinda, et al.
Veröffentlicht: (2025)
von: Jia, Xinda, et al.
Veröffentlicht: (2025)
Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
Inference Stage Denoising for Undersampled MRI Reconstruction
von: Xue, Yuyang, et al.
Veröffentlicht: (2024)
von: Xue, Yuyang, et al.
Veröffentlicht: (2024)
TwiSTAR:Think Fast, Think Slow, Then Act,Generative Recommendation with Adaptive Reasoning
von: Cao, Shiteng, et al.
Veröffentlicht: (2026)
von: Cao, Shiteng, et al.
Veröffentlicht: (2026)
From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models
von: Shen, Yi, et al.
Veröffentlicht: (2025)
von: Shen, Yi, et al.
Veröffentlicht: (2025)
When in Doubt, Think Slow: Iterative Reasoning with Latent Imagination
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2024)
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2024)
Agents Thinking Fast and Slow: A Talker-Reasoner Architecture
von: Christakopoulou, Konstantina, et al.
Veröffentlicht: (2024)
von: Christakopoulou, Konstantina, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
von: Li, Junsong, et al.
Veröffentlicht: (2025) -
Reinforced Interactive Continual Learning via Real-time Noisy Human Feedback
von: Yang, Yutao, et al.
Veröffentlicht: (2025) -
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
von: Li, Junsong, et al.
Veröffentlicht: (2025) -
Black-box Model Merging for Language-Model-as-a-Service with Massive Model Repositories
von: Chen, Shilian, et al.
Veröffentlicht: (2025) -
Forget What's Sensitive, Remember What Matters: Token-Level Differential Privacy in Memory Sculpting for Continual Learning
von: Zhan, Bihao, et al.
Veröffentlicht: (2025)