RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xie, Yu, Ren, Xingkai, Qi, Ying, Hu, Yao, Shan, Lianlei |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
par: Ye, Wengao, et autres
Publié: (2025)
par: Ye, Wengao, et autres
Publié: (2025)
Latent-Space Contrastive Reinforcement Learning for Stable and Efficient LLM Reasoning
par: Shan, Lianlei, et autres
Publié: (2026)
par: Shan, Lianlei, et autres
Publié: (2026)
GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models
par: Yi, Qiang, et autres
Publié: (2025)
par: Yi, Qiang, et autres
Publié: (2025)
Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning
par: Liu, Dong, et autres
Publié: (2026)
par: Liu, Dong, et autres
Publié: (2026)
Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs
par: Xia, Yu, et autres
Publié: (2024)
par: Xia, Yu, et autres
Publié: (2024)
RecGPT: Generative Personalized Prompts for Sequential Recommendation via ChatGPT Training Paradigm
par: Zhang, Yabin, et autres
Publié: (2024)
par: Zhang, Yabin, et autres
Publié: (2024)
Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
par: Chen, Lei, et autres
Publié: (2025)
par: Chen, Lei, et autres
Publié: (2025)
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
par: Shen, Maohao, et autres
Publié: (2025)
par: Shen, Maohao, et autres
Publié: (2025)
Pentest-R1: Towards Autonomous Penetration Testing Reasoning Optimized via Two-Stage Reinforcement Learning
par: Kong, He, et autres
Publié: (2025)
par: Kong, He, et autres
Publié: (2025)
Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning
par: Chen, Yiqun, et autres
Publié: (2026)
par: Chen, Yiqun, et autres
Publié: (2026)
ThinkDrive: Chain-of-Thought Guided Progressive Reinforcement Learning Fine-Tuning for Autonomous Driving
par: Zhao, Chang, et autres
Publié: (2026)
par: Zhao, Chang, et autres
Publié: (2026)
Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning
par: Xi, Zhiheng, et autres
Publié: (2025)
par: Xi, Zhiheng, et autres
Publié: (2025)
NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding
par: Liu, Shiyu, et autres
Publié: (2025)
par: Liu, Shiyu, et autres
Publié: (2025)
Lifelong Learning and Selective Forgetting via Contrastive Strategy
par: Shan, Lianlei, et autres
Publié: (2024)
par: Shan, Lianlei, et autres
Publié: (2024)
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
par: Feng, Zhaopeng, et autres
Publié: (2025)
par: Feng, Zhaopeng, et autres
Publié: (2025)
In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought
par: Huang, Sili, et autres
Publié: (2024)
par: Huang, Sili, et autres
Publié: (2024)
MuonRec: Shifting the Optimizer Paradigm Beyond Adam in Scalable Generative Recommendation
par: Shan, Rong, et autres
Publié: (2026)
par: Shan, Rong, et autres
Publié: (2026)
KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
par: Xu, Wanshun, et autres
Publié: (2025)
par: Xu, Wanshun, et autres
Publié: (2025)
SAGE: Sequence-level Adaptive Gradient Evolution for Generative Recommendation
par: Xie, Yu, et autres
Publié: (2026)
par: Xie, Yu, et autres
Publié: (2026)
LLM Reasoning Is Latent, Not the Chain of Thought
par: Wang, Wenshuo
Publié: (2026)
par: Wang, Wenshuo
Publié: (2026)
RecBundle: A Next-Generation Geometric Paradigm for Explainable Recommender Systems
par: Wang, Hui, et autres
Publié: (2026)
par: Wang, Hui, et autres
Publié: (2026)
Training Multimodal Large Reasoning Models Needs Better Thoughts: A Three-Stage Framework for Long Chain-of-Thought Synthesis and Selection
par: Wang, Yizhi, et autres
Publié: (2025)
par: Wang, Yizhi, et autres
Publié: (2025)
Learning to Correct: Calibrated Reinforcement Learning for Multi-Attempt Chain-of-Thought
par: Ildiz, Muhammed Emrullah, et autres
Publié: (2026)
par: Ildiz, Muhammed Emrullah, et autres
Publié: (2026)
Rethinking Supply Chain Planning: A Generative Paradigm
par: Yin, Jiaheng, et autres
Publié: (2025)
par: Yin, Jiaheng, et autres
Publié: (2025)
GOT4Rec: Graph of Thoughts for Sequential Recommendation
par: Long, Zewen, et autres
Publié: (2024)
par: Long, Zewen, et autres
Publié: (2024)
IDGenRec: LLM-RecSys Alignment with Textual ID Learning
par: Tan, Juntao, et autres
Publié: (2024)
par: Tan, Juntao, et autres
Publié: (2024)
CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning
par: Zhu, Wenqiao, et autres
Publié: (2025)
par: Zhu, Wenqiao, et autres
Publié: (2025)
Content-Aware Ad Banner Layout Generation with Two-Stage Chain-of-Thought in Vision Language Models
par: Yoshitake, Kei, et autres
Publié: (2025)
par: Yoshitake, Kei, et autres
Publié: (2025)
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
par: Hase, Peter, et autres
Publié: (2026)
par: Hase, Peter, et autres
Publié: (2026)
Cognitive Memory in Large Language Models
par: Shan, Lianlei, et autres
Publié: (2025)
par: Shan, Lianlei, et autres
Publié: (2025)
GCoT: Chain-of-Thought Prompt Learning for Graphs
par: Yu, Xingtong, et autres
Publié: (2025)
par: Yu, Xingtong, et autres
Publié: (2025)
Reinforcing Structured Chain-of-Thought for Video Understanding
par: Wang, Peiyao, et autres
Publié: (2026)
par: Wang, Peiyao, et autres
Publié: (2026)
Reinforcement Learning for Chain of Thought Compression with One-Domain-to-All Generalization
par: Li, Hanyu, et autres
Publié: (2025)
par: Li, Hanyu, et autres
Publié: (2025)
Empathy-R1: A Chain-of-Empathy and Reinforcement Learning Framework for Long-Form Mental Health Support
par: Yao, Xianrong, et autres
Publié: (2025)
par: Yao, Xianrong, et autres
Publié: (2025)
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
par: Zhang, Zhenguo, et autres
Publié: (2025)
par: Zhang, Zhenguo, et autres
Publié: (2025)
Learning Composable Chains-of-Thought
par: Yin, Fangcong, et autres
Publié: (2025)
par: Yin, Fangcong, et autres
Publié: (2025)
Reason from Future: Reverse Thought Chain Enhances LLM Reasoning
par: Xu, Yinlong, et autres
Publié: (2025)
par: Xu, Yinlong, et autres
Publié: (2025)
RationAnomaly: Log Anomaly Detection with Rationality via Chain-of-Thought and Reinforcement Learning
par: Xu, Song, et autres
Publié: (2025)
par: Xu, Song, et autres
Publié: (2025)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
par: Sheng, Leheng, et autres
Publié: (2026)
par: Sheng, Leheng, et autres
Publié: (2026)
RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
par: Hao, Qianyue, et autres
Publié: (2025)
par: Hao, Qianyue, et autres
Publié: (2025)
Documents similaires
-
Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
par: Ye, Wengao, et autres
Publié: (2025) -
Latent-Space Contrastive Reinforcement Learning for Stable and Efficient LLM Reasoning
par: Shan, Lianlei, et autres
Publié: (2026) -
GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models
par: Yi, Qiang, et autres
Publié: (2025) -
Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning
par: Liu, Dong, et autres
Publié: (2026) -
Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs
par: Xia, Yu, et autres
Publié: (2024)