Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Jali, Neharika, Nayak, Anupam, Joshi, Gauri |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
by: Patel, Shivam, et al.
Published: (2025)
by: Patel, Shivam, et al.
Published: (2025)
MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models
by: Raje, Arian, et al.
Published: (2026)
by: Raje, Arian, et al.
Published: (2026)
Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems
by: Jali, Neharika, et al.
Published: (2024)
by: Jali, Neharika, et al.
Published: (2024)
Erasure Coded Neural Network Inference via Fisher Averaging
by: Jhunjhunwala, Divyansh, et al.
Published: (2024)
by: Jhunjhunwala, Divyansh, et al.
Published: (2024)
Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking
by: He, Zhida, et al.
Published: (2026)
by: He, Zhida, et al.
Published: (2026)
Natural Policy Gradient for Average Reward Non-Stationary RL
by: Jali, Neharika, et al.
Published: (2025)
by: Jali, Neharika, et al.
Published: (2025)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
by: Ding, Yifeng, et al.
Published: (2025)
by: Ding, Yifeng, et al.
Published: (2025)
ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical Dialogue
by: Cao, Ruike, et al.
Published: (2026)
by: Cao, Ruike, et al.
Published: (2026)
CARE: Turning LLMs Into Causal Reasoning Expert
by: Dong, Juncheng, et al.
Published: (2025)
by: Dong, Juncheng, et al.
Published: (2025)
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
by: Askin, Baris, et al.
Published: (2026)
by: Askin, Baris, et al.
Published: (2026)
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
by: Jeffares, Alan, et al.
Published: (2025)
by: Jeffares, Alan, et al.
Published: (2025)
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
by: Liu, Licheng, et al.
Published: (2025)
by: Liu, Licheng, et al.
Published: (2025)
Didactic to Constructive: Turning Expert Solutions into Learnable Reasoning
by: Mendes, Ethan, et al.
Published: (2026)
by: Mendes, Ethan, et al.
Published: (2026)
LOCUS: Low-Dimensional Model Embeddings for Efficient Model Exploration, Comparison, and Selection
by: Patel, Shivam, et al.
Published: (2026)
by: Patel, Shivam, et al.
Published: (2026)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
by: Goru, Ritesh, et al.
Published: (2025)
by: Goru, Ritesh, et al.
Published: (2025)
Mitigating Conversational Inertia in Multi-Turn Agents
by: Wan, Yang, et al.
Published: (2026)
by: Wan, Yang, et al.
Published: (2026)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
Not All Instances Are Equally Valuable: Towards Influence-Weighted Dataset Distillation
by: Deng, Qiyan, et al.
Published: (2025)
by: Deng, Qiyan, et al.
Published: (2025)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
by: Kumarappan, Adarsh, et al.
Published: (2025)
by: Kumarappan, Adarsh, et al.
Published: (2025)
Conformal Thinking: Risk Control for Reasoning on a Compute Budget
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Ravan: Multi-Head Low-Rank Adaptation for Federated Fine-Tuning
by: Raje, Arian, et al.
Published: (2025)
by: Raje, Arian, et al.
Published: (2025)
Not All Timesteps Matter Equally: Selective Alignment Knowledge Distillation for Spiking Neural Networks
by: Sun, Kai, et al.
Published: (2026)
by: Sun, Kai, et al.
Published: (2026)
Hard Examples Are All You Need: Maximizing GRPO Post-Training Under Annotation Budgets
by: Pikus, Benjamin, et al.
Published: (2025)
by: Pikus, Benjamin, et al.
Published: (2025)
Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation
by: Wei, Chenxing, et al.
Published: (2026)
by: Wei, Chenxing, et al.
Published: (2026)
POLO: Preference-Guided Multi-Turn Reinforcement Learning for Lead Optimization
by: Wang, Ziqing, et al.
Published: (2025)
by: Wang, Ziqing, et al.
Published: (2025)
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
by: Xiang, Violet, et al.
Published: (2025)
by: Xiang, Violet, et al.
Published: (2025)
Investigating the Treacherous Turn in Deep Reinforcement Learning
by: Ashcraft, Chace, et al.
Published: (2025)
by: Ashcraft, Chace, et al.
Published: (2025)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
by: Rahman, Salman, et al.
Published: (2025)
by: Rahman, Salman, et al.
Published: (2025)
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory
by: Li, Sijia, et al.
Published: (2025)
by: Li, Sijia, et al.
Published: (2025)
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
by: Ekbote, Chanakya, et al.
Published: (2025)
by: Ekbote, Chanakya, et al.
Published: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
On Arbitrary Predictions from Equally Valid Models
by: Lockfisch, Sarah, et al.
Published: (2025)
by: Lockfisch, Sarah, et al.
Published: (2025)
Equally Critical: Samples, Targets, and Their Mappings in Datasets
by: Yang, Runkang, et al.
Published: (2025)
by: Yang, Runkang, et al.
Published: (2025)
Multi-Turn Code Generation Through Single-Step Rewards
by: Jain, Arnav Kumar, et al.
Published: (2025)
by: Jain, Arnav Kumar, et al.
Published: (2025)
Turn-based Multi-Agent Reinforcement Learning Model Checking
by: Gross, Dennis
Published: (2025)
by: Gross, Dennis
Published: (2025)
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
Indirect Attention: Turning Context Misalignment into a Feature
by: Bahaduri, Bissmella, et al.
Published: (2025)
by: Bahaduri, Bissmella, et al.
Published: (2025)
Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
by: Yang, Xiaoyu, et al.
Published: (2025)
by: Yang, Xiaoyu, et al.
Published: (2025)
Kevin: Multi-Turn RL for Generating CUDA Kernels
by: Baronio, Carlo, et al.
Published: (2025)
by: Baronio, Carlo, et al.
Published: (2025)
Think in Blocks: Adaptive Reasoning from Direct Response to Deep Reasoning
by: Zhu, Yekun, et al.
Published: (2025)
by: Zhu, Yekun, et al.
Published: (2025)
Similar Items
-
ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
by: Patel, Shivam, et al.
Published: (2025) -
MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models
by: Raje, Arian, et al.
Published: (2026) -
Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems
by: Jali, Neharika, et al.
Published: (2024) -
Erasure Coded Neural Network Inference via Fisher Averaging
by: Jhunjhunwala, Divyansh, et al.
Published: (2024) -
Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking
by: He, Zhida, et al.
Published: (2026)