Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Ziyan, Wang, Zheng, Qu, Xingwei, Cheng, Qi, Fu, Jie, Tang, Shengpu, Zhang, Minjia, Huo, Xiaoming |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
di: Wang, Ziyan, et al.
Pubblicazione: (2025)
di: Wang, Ziyan, et al.
Pubblicazione: (2025)
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
di: Paliotta, Daniele, et al.
Pubblicazione: (2025)
di: Paliotta, Daniele, et al.
Pubblicazione: (2025)
Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
di: Xu, Xu, et al.
Pubblicazione: (2025)
di: Xu, Xu, et al.
Pubblicazione: (2025)
Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
di: Min, Yingqian, et al.
Pubblicazione: (2024)
di: Min, Yingqian, et al.
Pubblicazione: (2024)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
di: Qi, Biqing, et al.
Pubblicazione: (2024)
di: Qi, Biqing, et al.
Pubblicazione: (2024)
Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents
di: Yang, Ruihan, et al.
Pubblicazione: (2026)
di: Yang, Ruihan, et al.
Pubblicazione: (2026)
Thinker: Learning to Think Fast and Slow
di: Chung, Stephen, et al.
Pubblicazione: (2025)
di: Chung, Stephen, et al.
Pubblicazione: (2025)
Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
di: Du, Y., et al.
Pubblicazione: (2025)
di: Du, Y., et al.
Pubblicazione: (2025)
Agents Thinking Fast and Slow: A Talker-Reasoner Architecture
di: Christakopoulou, Konstantina, et al.
Pubblicazione: (2024)
di: Christakopoulou, Konstantina, et al.
Pubblicazione: (2024)
CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
di: Tang, Zhengyang, et al.
Pubblicazione: (2025)
di: Tang, Zhengyang, et al.
Pubblicazione: (2025)
Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing
di: Yuan, Wenhao, et al.
Pubblicazione: (2026)
di: Yuan, Wenhao, et al.
Pubblicazione: (2026)
Micro-Act: Mitigating Knowledge Conflict in LLM-based RAG via Actionable Self-Reasoning
di: Huo, Nan, et al.
Pubblicazione: (2025)
di: Huo, Nan, et al.
Pubblicazione: (2025)
Look Before You Leap: Autonomous Exploration for LLM Agents
di: Ye, Ziang, et al.
Pubblicazione: (2026)
di: Ye, Ziang, et al.
Pubblicazione: (2026)
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation
di: Zheng, Tianyu, et al.
Pubblicazione: (2024)
di: Zheng, Tianyu, et al.
Pubblicazione: (2024)
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
di: He, Yanjie
Pubblicazione: (2026)
di: He, Yanjie
Pubblicazione: (2026)
CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor Generation
di: Shao, Yujie, et al.
Pubblicazione: (2024)
di: Shao, Yujie, et al.
Pubblicazione: (2024)
Enhancing LLM Reasoning with Reward-guided Tree Search
di: Jiang, Jinhao, et al.
Pubblicazione: (2024)
di: Jiang, Jinhao, et al.
Pubblicazione: (2024)
Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
di: Lu, Junjie, et al.
Pubblicazione: (2025)
di: Lu, Junjie, et al.
Pubblicazione: (2025)
Learning to Reason under Off-Policy Guidance
di: Yan, Jianhao, et al.
Pubblicazione: (2025)
di: Yan, Jianhao, et al.
Pubblicazione: (2025)
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
di: Fu, Tingchen, et al.
Pubblicazione: (2025)
di: Fu, Tingchen, et al.
Pubblicazione: (2025)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
di: Hong, Yining, et al.
Pubblicazione: (2024)
di: Hong, Yining, et al.
Pubblicazione: (2024)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
di: Xiao, Wenyi, et al.
Pubblicazione: (2025)
di: Xiao, Wenyi, et al.
Pubblicazione: (2025)
Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden Principles
di: Wei, Qingyan, et al.
Pubblicazione: (2025)
di: Wei, Qingyan, et al.
Pubblicazione: (2025)
Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents
di: Zeng, Jingying, et al.
Pubblicazione: (2025)
di: Zeng, Jingying, et al.
Pubblicazione: (2025)
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
di: Qi, Penghui, et al.
Pubblicazione: (2025)
di: Qi, Penghui, et al.
Pubblicazione: (2025)
Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning
di: He, Jiashu, et al.
Pubblicazione: (2026)
di: He, Jiashu, et al.
Pubblicazione: (2026)
Optimizing Length Compression in Large Reasoning Models
di: Cheng, Zhengxiang, et al.
Pubblicazione: (2025)
di: Cheng, Zhengxiang, et al.
Pubblicazione: (2025)
ExperienceWeaver: Optimizing Small-sample Experience Learning for LLM-based Clinical Text Improvement
di: Xiao, Ziyan, et al.
Pubblicazione: (2026)
di: Xiao, Ziyan, et al.
Pubblicazione: (2026)
A Neurosymbolic Fast and Slow Architecture for Graph Coloring
di: Khandelwal, Vedant, et al.
Pubblicazione: (2024)
di: Khandelwal, Vedant, et al.
Pubblicazione: (2024)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
di: Lu, Jinghui, et al.
Pubblicazione: (2025)
di: Lu, Jinghui, et al.
Pubblicazione: (2025)
Alleviating Choice Supportive Bias in LLM with Reasoning Dependency Generation
di: Zhuang, Nan, et al.
Pubblicazione: (2025)
di: Zhuang, Nan, et al.
Pubblicazione: (2025)
Towards Domain Adaptive Neural Contextual Bandits
di: Wang, Ziyan, et al.
Pubblicazione: (2024)
di: Wang, Ziyan, et al.
Pubblicazione: (2024)
Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization
di: Simoni, Marco, et al.
Pubblicazione: (2025)
di: Simoni, Marco, et al.
Pubblicazione: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
di: Zhang, Xichen, et al.
Pubblicazione: (2025)
di: Zhang, Xichen, et al.
Pubblicazione: (2025)
The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues
di: Cheng, Youyou, et al.
Pubblicazione: (2026)
di: Cheng, Youyou, et al.
Pubblicazione: (2026)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
di: Li, Changming, et al.
Pubblicazione: (2026)
di: Li, Changming, et al.
Pubblicazione: (2026)
Read Before You Think: Mitigating LLM Comprehension Failures with Step-by-Step Reading
di: Han, Feijiang, et al.
Pubblicazione: (2025)
di: Han, Feijiang, et al.
Pubblicazione: (2025)
Hierarchical Budget Policy Optimization for Adaptive Reasoning
di: Lyu, Shangke, et al.
Pubblicazione: (2025)
di: Lyu, Shangke, et al.
Pubblicazione: (2025)
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
di: Zhang, Zhiwei, et al.
Pubblicazione: (2025)
di: Zhang, Zhiwei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
di: Wang, Ziyan, et al.
Pubblicazione: (2025) -
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
di: Paliotta, Daniele, et al.
Pubblicazione: (2025) -
Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
di: Xu, Xu, et al.
Pubblicazione: (2025) -
Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
di: Min, Yingqian, et al.
Pubblicazione: (2024) -
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
di: Qi, Biqing, et al.
Pubblicazione: (2024)