CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Gan, Zeyu, Yi, Hao, Liu, Yong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning
por: Gan, Zeyu, et al.
Publicado: (2025)
por: Gan, Zeyu, et al.
Publicado: (2025)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
por: Deng, Yuntian, et al.
Publicado: (2024)
por: Deng, Yuntian, et al.
Publicado: (2024)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
por: Yao, Xinhao, et al.
Publicado: (2025)
por: Yao, Xinhao, et al.
Publicado: (2025)
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
por: Gan, Zeyu, et al.
Publicado: (2024)
por: Gan, Zeyu, et al.
Publicado: (2024)
CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
por: Bi, Jinhe, et al.
Publicado: (2025)
por: Bi, Jinhe, et al.
Publicado: (2025)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
por: Lee, Seongyun, et al.
Publicado: (2025)
por: Lee, Seongyun, et al.
Publicado: (2025)
Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
por: Moore, Kyle, et al.
Publicado: (2024)
por: Moore, Kyle, et al.
Publicado: (2024)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
por: Ren, Ruifeng, et al.
Publicado: (2024)
por: Ren, Ruifeng, et al.
Publicado: (2024)
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
por: Xu, Haotian, et al.
Publicado: (2025)
por: Xu, Haotian, et al.
Publicado: (2025)
MyGO Multiplex CoT: A Method for Self-Reflection in Large Language Models via Double Chain of Thought Thinking
por: Ji, Shihao, et al.
Publicado: (2025)
por: Ji, Shihao, et al.
Publicado: (2025)
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
por: Ye, Xinwu, et al.
Publicado: (2026)
por: Ye, Xinwu, et al.
Publicado: (2026)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
por: Xiao, Wenyi, et al.
Publicado: (2025)
por: Xiao, Wenyi, et al.
Publicado: (2025)
SIM-CoT: Supervised Implicit Chain-of-Thought
por: Wei, Xilin, et al.
Publicado: (2025)
por: Wei, Xilin, et al.
Publicado: (2025)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
por: Sprague, Zayne, et al.
Publicado: (2024)
por: Sprague, Zayne, et al.
Publicado: (2024)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
por: Zhou, Weibo, et al.
Publicado: (2025)
por: Zhou, Weibo, et al.
Publicado: (2025)
Efficient Long CoT Reasoning in Small Language Models
por: Wang, Zhaoyang, et al.
Publicado: (2025)
por: Wang, Zhaoyang, et al.
Publicado: (2025)
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
por: Luo, Haotian, et al.
Publicado: (2025)
por: Luo, Haotian, et al.
Publicado: (2025)
Effectiveness of Zero-shot-CoT in Japanese Prompts
por: Takayama, Shusuke, et al.
Publicado: (2025)
por: Takayama, Shusuke, et al.
Publicado: (2025)
CoT-Driven Framework for Short Text Classification: Enhancing and Transferring Capabilities from Large to Smaller Model
por: Wu, Hui, et al.
Publicado: (2024)
por: Wu, Hui, et al.
Publicado: (2024)
CoT-Valve: Length-Compressible Chain-of-Thought Tuning
por: Ma, Xinyin, et al.
Publicado: (2025)
por: Ma, Xinyin, et al.
Publicado: (2025)
Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens
por: Yong, Xixian, et al.
Publicado: (2025)
por: Yong, Xixian, et al.
Publicado: (2025)
TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models
por: Gao, Fan, et al.
Publicado: (2025)
por: Gao, Fan, et al.
Publicado: (2025)
CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context Demonstrations
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework
por: Cui, Jin, et al.
Publicado: (2026)
por: Cui, Jin, et al.
Publicado: (2026)
CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
por: Dong, Qihua, et al.
Publicado: (2025)
por: Dong, Qihua, et al.
Publicado: (2025)
KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
por: Mondal, Debjyoti, et al.
Publicado: (2024)
por: Mondal, Debjyoti, et al.
Publicado: (2024)
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
por: Li, Jiachun, et al.
Publicado: (2024)
por: Li, Jiachun, et al.
Publicado: (2024)
Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
por: Du, Y., et al.
Publicado: (2025)
por: Du, Y., et al.
Publicado: (2025)
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
por: Bilal, Ahsan, et al.
Publicado: (2025)
por: Bilal, Ahsan, et al.
Publicado: (2025)
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
por: Paliotta, Daniele, et al.
Publicado: (2025)
por: Paliotta, Daniele, et al.
Publicado: (2025)
MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
por: Kim, Soo Yong, et al.
Publicado: (2025)
por: Kim, Soo Yong, et al.
Publicado: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
por: Guo, Ziyu, et al.
Publicado: (2025)
por: Guo, Ziyu, et al.
Publicado: (2025)
Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs
por: Chatterjee, Sagnik, et al.
Publicado: (2026)
por: Chatterjee, Sagnik, et al.
Publicado: (2026)
CoT-BERT: Enhancing Unsupervised Sentence Representation through Chain-of-Thought
por: Zhang, Bowen, et al.
Publicado: (2023)
por: Zhang, Bowen, et al.
Publicado: (2023)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
por: Yang, Cehao, et al.
Publicado: (2025)
por: Yang, Cehao, et al.
Publicado: (2025)
Fast-Slow-Thinking: Complex Task Solving with Large Language Models
por: Sun, Yiliu, et al.
Publicado: (2025)
por: Sun, Yiliu, et al.
Publicado: (2025)
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
por: Cai, Wenrui, et al.
Publicado: (2025)
por: Cai, Wenrui, et al.
Publicado: (2025)
CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
por: Shao, Jintian, et al.
Publicado: (2025)
por: Shao, Jintian, et al.
Publicado: (2025)
CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor Generation
por: Shao, Yujie, et al.
Publicado: (2024)
por: Shao, Yujie, et al.
Publicado: (2024)
Ejemplares similares
-
Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning
por: Gan, Zeyu, et al.
Publicado: (2025) -
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
por: Deng, Yuntian, et al.
Publicado: (2024) -
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
por: Yao, Xinhao, et al.
Publicado: (2025) -
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
por: Gan, Zeyu, et al.
Publicado: (2024) -
CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
por: Bi, Jinhe, et al.
Publicado: (2025)