When Actions Teach You to Think: Reasoning-Action Synergy via Reinforcement Learning in Conversational Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rawat, Mrinal, Chakraborty, Arkajyoti, Gupta, Neha, Pieraccini, Roberto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
REFINE on Scarce Data: Retrieval Enhancement through Fine-Tuning via Model Fusion of Embedding Models
von: Gupta, Ambuje, et al.
Veröffentlicht: (2024)
von: Gupta, Ambuje, et al.
Veröffentlicht: (2024)
Mechanistic Anomaly Detection for "Quirky" Language Models
von: Johnston, David O., et al.
Veröffentlicht: (2025)
von: Johnston, David O., et al.
Veröffentlicht: (2025)
Controllable Discovery of Intents: Incremental Deep Clustering Using Semi-Supervised Contrastive Learning
von: Rawat, Mrinal, et al.
Veröffentlicht: (2024)
von: Rawat, Mrinal, et al.
Veröffentlicht: (2024)
AdaptThink: Reasoning Models Can Learn When to Think
von: Zhang, Jiajie, et al.
Veröffentlicht: (2025)
von: Zhang, Jiajie, et al.
Veröffentlicht: (2025)
Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions
von: Li, Yongqi, et al.
Veröffentlicht: (2026)
von: Li, Yongqi, et al.
Veröffentlicht: (2026)
Think When You Need: Self-Adaptive Chain-of-Thought Learning
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
LiveThinking: Enabling Real-Time Efficient Reasoning for AI-Powered Livestreaming via Reinforcement Learning
von: Sun, Yuhan, et al.
Veröffentlicht: (2025)
von: Sun, Yuhan, et al.
Veröffentlicht: (2025)
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
von: Zhao, Siyan, et al.
Veröffentlicht: (2025)
von: Zhao, Siyan, et al.
Veröffentlicht: (2025)
More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents
von: Xiong, Weimin, et al.
Veröffentlicht: (2025)
von: Xiong, Weimin, et al.
Veröffentlicht: (2025)
Learning When to Quit in Sales Conversations
von: Manzoor, Emaad, et al.
Veröffentlicht: (2025)
von: Manzoor, Emaad, et al.
Veröffentlicht: (2025)
Internalizing LLM Reasoning via Discovery and Replay of Latent Actions
von: Shi, Zhenning, et al.
Veröffentlicht: (2026)
von: Shi, Zhenning, et al.
Veröffentlicht: (2026)
Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code
von: Bao, Keqin, et al.
Veröffentlicht: (2025)
von: Bao, Keqin, et al.
Veröffentlicht: (2025)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
Think Before You Lie: How Reasoning Leads to Honesty
von: Yuan, Ann, et al.
Veröffentlicht: (2026)
von: Yuan, Ann, et al.
Veröffentlicht: (2026)
Pre-Act: Multi-Step Planning and Reasoning Improves Acting in LLM Agents
von: Rawat, Mrinal, et al.
Veröffentlicht: (2025)
von: Rawat, Mrinal, et al.
Veröffentlicht: (2025)
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
Reinforcement Learning is all You Need
von: Lian, Yongsheng
Veröffentlicht: (2025)
von: Lian, Yongsheng
Veröffentlicht: (2025)
DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
Reasoning with Sampling: Your Base Model is Smarter Than You Think
von: Karan, Aayush, et al.
Veröffentlicht: (2025)
von: Karan, Aayush, et al.
Veröffentlicht: (2025)
When Less is Enough: Efficient Inference via Collaborative Reasoning
von: Chen, Yilei, et al.
Veröffentlicht: (2026)
von: Chen, Yilei, et al.
Veröffentlicht: (2026)
The RL/LLM Taxonomy Tree: Reviewing Synergies Between Reinforcement Learning and Large Language Models
von: Pternea, Moschoula, et al.
Veröffentlicht: (2024)
von: Pternea, Moschoula, et al.
Veröffentlicht: (2024)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
Teaching Language Models to Critique via Reinforcement Learning
von: Xie, Zhihui, et al.
Veröffentlicht: (2025)
von: Xie, Zhihui, et al.
Veröffentlicht: (2025)
S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
von: Ma, Ruotian, et al.
Veröffentlicht: (2025)
von: Ma, Ruotian, et al.
Veröffentlicht: (2025)
Language Model Cascades: Token-level uncertainty and beyond
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
von: Yang, Rui, et al.
Veröffentlicht: (2026)
von: Yang, Rui, et al.
Veröffentlicht: (2026)
Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency
von: Lin, Victoria, et al.
Veröffentlicht: (2026)
von: Lin, Victoria, et al.
Veröffentlicht: (2026)
Agents Thinking Fast and Slow: A Talker-Reasoner Architecture
von: Christakopoulou, Konstantina, et al.
Veröffentlicht: (2024)
von: Christakopoulou, Konstantina, et al.
Veröffentlicht: (2024)
CaRT: Teaching LLM Agents to Know When They Know Enough
von: Liu, Grace, et al.
Veröffentlicht: (2025)
von: Liu, Grace, et al.
Veröffentlicht: (2025)
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
von: Sundaram, Shobhita, et al.
Veröffentlicht: (2026)
von: Sundaram, Shobhita, et al.
Veröffentlicht: (2026)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Tool-Integrated Interleaved Thinking towards Cross-Domain Generalization
von: Chen, Zhengyu, et al.
Veröffentlicht: (2025)
von: Chen, Zhengyu, et al.
Veröffentlicht: (2025)
ExpThink: Experience-Guided Reinforcement Learning for Adaptive Chain-of-Thought Compression
von: Bian, Tingcheng, et al.
Veröffentlicht: (2026)
von: Bian, Tingcheng, et al.
Veröffentlicht: (2026)
HierRouter: Coordinated Routing of Specialized Large Language Models via Reinforcement Learning
von: Gupta, Nikunj, et al.
Veröffentlicht: (2025)
von: Gupta, Nikunj, et al.
Veröffentlicht: (2025)
When Domains Interact: Asymmetric and Order-Sensitive Cross-Domain Effects in Reinforcement Learning for Reasoning
von: Yang, Wang, et al.
Veröffentlicht: (2026)
von: Yang, Wang, et al.
Veröffentlicht: (2026)
LoRA Is Slower Than You Think
von: Ko, Seokmin
Veröffentlicht: (2025)
von: Ko, Seokmin
Veröffentlicht: (2025)
ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models
von: Sun, Chung-En, et al.
Veröffentlicht: (2025)
von: Sun, Chung-En, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
REFINE on Scarce Data: Retrieval Enhancement through Fine-Tuning via Model Fusion of Embedding Models
von: Gupta, Ambuje, et al.
Veröffentlicht: (2024) -
Mechanistic Anomaly Detection for "Quirky" Language Models
von: Johnston, David O., et al.
Veröffentlicht: (2025) -
Controllable Discovery of Intents: Incremental Deep Clustering Using Semi-Supervised Contrastive Learning
von: Rawat, Mrinal, et al.
Veröffentlicht: (2024) -
AdaptThink: Reasoning Models Can Learn When to Think
von: Zhang, Jiajie, et al.
Veröffentlicht: (2025) -
Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions
von: Li, Yongqi, et al.
Veröffentlicht: (2026)