Good Reasoning Makes Good Demonstrations: Implicit Reasoning Quality Supervision via In-Context Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mei, Tiehua, Lv, Minxuan, Pan, Leiyu, Su, Zhenpeng, Hou, Hongru, Chen, Hengrui, Xu, Ao, Yang, Deqing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
von: Hou, Hongru, et al.
Veröffentlicht: (2026)
von: Hou, Hongru, et al.
Veröffentlicht: (2026)
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
von: Lv, Minxuan, et al.
Veröffentlicht: (2026)
von: Lv, Minxuan, et al.
Veröffentlicht: (2026)
GORACS: Group-level Optimal Transport-guided Coreset Selection for LLM-based Recommender Systems
von: Mei, Tiehua, et al.
Veröffentlicht: (2025)
von: Mei, Tiehua, et al.
Veröffentlicht: (2025)
Guiding Through Complexity: What Makes Good Supervision for Hard Math Reasoning Tasks?
von: He, Xuan, et al.
Veröffentlicht: (2024)
von: He, Xuan, et al.
Veröffentlicht: (2024)
Implicit Reasoning in Transformers is Reasoning through Shortcuts
von: Lin, Tianhe, et al.
Veröffentlicht: (2025)
von: Lin, Tianhe, et al.
Veröffentlicht: (2025)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
von: Ki, Dayeon, et al.
Veröffentlicht: (2026)
von: Ki, Dayeon, et al.
Veröffentlicht: (2026)
Hard Proofs and Good Reasons
von: DeDeo, Simon
Veröffentlicht: (2024)
von: DeDeo, Simon
Veröffentlicht: (2024)
Good Data Makes Good Sense.
von: Anderson, Mary Alice
Veröffentlicht: (1999)
von: Anderson, Mary Alice
Veröffentlicht: (1999)
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
von: Cheng, Junhao, et al.
Veröffentlicht: (2026)
von: Cheng, Junhao, et al.
Veröffentlicht: (2026)
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
von: Jiang, Gangwei, et al.
Veröffentlicht: (2025)
von: Jiang, Gangwei, et al.
Veröffentlicht: (2025)
What Makes Good In-context Demonstrations for Code Intelligence Tasks with LLMs?
von: Gao, Shuzheng, et al.
Veröffentlicht: (2023)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2023)
Finedeep: Mitigating Sparse Activation in Dense LLMs via Multi-Layer Fine-Grained Experts
von: Pan, Leiyu, et al.
Veröffentlicht: (2025)
von: Pan, Leiyu, et al.
Veröffentlicht: (2025)
Nonlinearity in Dynamic Causal Effects: Making the Bad into the Good, and the Good into the Great?
von: Kitagawa, Toru, et al.
Veröffentlicht: (2025)
von: Kitagawa, Toru, et al.
Veröffentlicht: (2025)
What Makes Good Instruction-Tuning Data? An In-Context Learning Perspective
von: Han, Guangzeng, et al.
Veröffentlicht: (2026)
von: Han, Guangzeng, et al.
Veröffentlicht: (2026)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
Enhancing In-Context Learning via Implicit Demonstration Augmentation
von: Zhou, Xiaoling, et al.
Veröffentlicht: (2024)
von: Zhou, Xiaoling, et al.
Veröffentlicht: (2024)
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
What Makes a Good Curriculum? Disentangling the Effects of Data Ordering on LLM Mathematical Reasoning
von: Jia, Yaning, et al.
Veröffentlicht: (2025)
von: Jia, Yaning, et al.
Veröffentlicht: (2025)
DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMs
von: Lv, Minxuan, et al.
Veröffentlicht: (2025)
von: Lv, Minxuan, et al.
Veröffentlicht: (2025)
Learning to Detect Baked Goods with Limited Supervision
von: Schmitt, Thomas H., et al.
Veröffentlicht: (2026)
von: Schmitt, Thomas H., et al.
Veröffentlicht: (2026)
LAD-Reasoner: Tiny Multimodal Models are Good Reasoners for Logical Anomaly Detection
von: Li, Weijia, et al.
Veröffentlicht: (2025)
von: Li, Weijia, et al.
Veröffentlicht: (2025)
Good News Is Not a Sufficient Condition for Motivated Reasoning
von: Thaler, Michael
Veröffentlicht: (2020)
von: Thaler, Michael
Veröffentlicht: (2020)
Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis
von: Jin, Hongbo, et al.
Veröffentlicht: (2026)
von: Jin, Hongbo, et al.
Veröffentlicht: (2026)
Making a Good Doctor
Veröffentlicht: (2025)
Veröffentlicht: (2025)
Making Good in The Global South
von: Carolina Villagra
Veröffentlicht: (2026)
von: Carolina Villagra
Veröffentlicht: (2026)
Making the Library Good for Business
von: Lee, John W., et al.
Veröffentlicht: (1973)
von: Lee, John W., et al.
Veröffentlicht: (1973)
Multi-Path Collaborative Reasoning via Reinforcement Learning
von: Lv, Jindi, et al.
Veröffentlicht: (2025)
von: Lv, Jindi, et al.
Veröffentlicht: (2025)
Process In-Context Learning: Enhancing Mathematical Reasoning via Dynamic Demonstration Insertion
von: Gao, Ang, et al.
Veröffentlicht: (2026)
von: Gao, Ang, et al.
Veröffentlicht: (2026)
How do Transformers Learn Implicit Reasoning?
von: Ye, Jiaran, et al.
Veröffentlicht: (2025)
von: Ye, Jiaran, et al.
Veröffentlicht: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Good Seed Makes a Good Crop: Discovering Secret Seeds in Text-to-Image Diffusion Models
von: Xu, Katherine, et al.
Veröffentlicht: (2024)
von: Xu, Katherine, et al.
Veröffentlicht: (2024)
Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning
von: Hu, Wenbin, et al.
Veröffentlicht: (2025)
von: Hu, Wenbin, et al.
Veröffentlicht: (2025)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
von: Do, Heejin, et al.
Veröffentlicht: (2025)
von: Do, Heejin, et al.
Veröffentlicht: (2025)
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
von: Piedrahita, David Guzman, et al.
Veröffentlicht: (2025)
von: Piedrahita, David Guzman, et al.
Veröffentlicht: (2025)
Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
von: Zhou, Zihao, et al.
Veröffentlicht: (2024)
von: Zhou, Zihao, et al.
Veröffentlicht: (2024)
Demonstration Selection for In-Context Learning via Reinforcement Learning
von: Wang, Xubin, et al.
Veröffentlicht: (2024)
von: Wang, Xubin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
von: Hou, Hongru, et al.
Veröffentlicht: (2026) -
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025) -
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025) -
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025) -
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
von: Lv, Minxuan, et al.
Veröffentlicht: (2026)