Implicit Turn-Wise Policy Optimization for Proactive User-LLM Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Haoyu, Chen, Yuxin, Luo, Liang, Zhang, Buyun, Wen, Ellie Dingqiao, Li, Pan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Wukong: Towards a Scaling Law for Large-Scale Recommendation
by: Zhang, Buyun, et al.
Published: (2024)
by: Zhang, Buyun, et al.
Published: (2024)
Disaggregated Multi-Tower: Topology-aware Modeling Technique for Efficient Large-Scale Recommendation
by: Luo, Liang, et al.
Published: (2024)
by: Luo, Liang, et al.
Published: (2024)
Proactive Constrained Policy Optimization with Preemptive Penalty
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
by: Li, Jiaqian, et al.
Published: (2026)
by: Li, Jiaqian, et al.
Published: (2026)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Turning Stale Gradients into Stable Gradients: Coherent Coordinate Descent with Implicit Landscape Smoothing for Lightweight Zeroth-Order Optimization
by: Liang, Chen, et al.
Published: (2026)
by: Liang, Chen, et al.
Published: (2026)
Reward-Driven Interaction: Enhancing Proactive Dialogue Agents through User Satisfaction Prediction
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical Dialogue
by: Cao, Ruike, et al.
Published: (2026)
by: Cao, Ruike, et al.
Published: (2026)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
by: He, Junhui, et al.
Published: (2024)
by: He, Junhui, et al.
Published: (2024)
AlignIQL: Policy Alignment in Implicit Q-Learning through Constrained Optimization
by: He, Longxiang, et al.
Published: (2024)
by: He, Longxiang, et al.
Published: (2024)
The Efficiency vs. Accuracy Trade-off: Optimizing RAG-Enhanced LLM Recommender Systems Using Multi-Head Early Exit
by: Zhou, Huixue, et al.
Published: (2025)
by: Zhou, Huixue, et al.
Published: (2025)
DistDNAS: Search Efficient Feature Interactions within 2 Hours
by: Zhang, Tunhou, et al.
Published: (2023)
by: Zhang, Tunhou, et al.
Published: (2023)
Skip-Connected Policy Optimization for Implicit Advantage
by: Teng, Fengwei, et al.
Published: (2026)
by: Teng, Fengwei, et al.
Published: (2026)
Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants
by: Nathani, Deepak, et al.
Published: (2026)
by: Nathani, Deepak, et al.
Published: (2026)
Optimizing Personalized Federated Learning through Adaptive Layer-Wise Learning
by: Chen, Weihang, et al.
Published: (2024)
by: Chen, Weihang, et al.
Published: (2024)
Direct Regret Optimization in Bayesian Optimization
by: Zhang, Fengxue, et al.
Published: (2025)
by: Zhang, Fengxue, et al.
Published: (2025)
Extreme Value Policy Optimization for Safe Reinforcement Learning
by: Gao, Shiqing, et al.
Published: (2026)
by: Gao, Shiqing, et al.
Published: (2026)
IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
by: Luo, Haohao, et al.
Published: (2026)
by: Luo, Haohao, et al.
Published: (2026)
ILILT: Implicit Learning of Inverse Lithography Technologies
by: Yang, Haoyu, et al.
Published: (2024)
by: Yang, Haoyu, et al.
Published: (2024)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
by: Liang, Buyun, et al.
Published: (2025)
by: Liang, Buyun, et al.
Published: (2025)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
by: Ding, Yifeng, et al.
Published: (2025)
by: Ding, Yifeng, et al.
Published: (2025)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
by: Wu, Yuning, et al.
Published: (2026)
by: Wu, Yuning, et al.
Published: (2026)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Improving Efficiency of Iso-Surface Extraction on Implicit Neural Representations Using Uncertainty Propagation
by: Li, Haoyu, et al.
Published: (2024)
by: Li, Haoyu, et al.
Published: (2024)
Explorable INR: An Implicit Neural Representation for Ensemble Simulation Enabling Efficient Spatial and Parameter Exploration
by: Chen, Yi-Tang, et al.
Published: (2025)
by: Chen, Yi-Tang, et al.
Published: (2025)
DelayPTC-LLM: Metro Passenger Travel Choice Prediction under Train Delays with Large Language Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
On the Convergence of Single-Loop Stochastic Bilevel Optimization with Approximate Implicit Differentiation
by: Zhou, Yubo, et al.
Published: (2026)
by: Zhou, Yubo, et al.
Published: (2026)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization
by: Wang, Boxiao, et al.
Published: (2026)
by: Wang, Boxiao, et al.
Published: (2026)
POLO: Preference-Guided Multi-Turn Reinforcement Learning for Lead Optimization
by: Wang, Ziqing, et al.
Published: (2025)
by: Wang, Ziqing, et al.
Published: (2025)
Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
by: Hahn, Meera, et al.
Published: (2024)
by: Hahn, Meera, et al.
Published: (2024)
ISEP: Implicit Support Expansion for Offline Reinforcement Learning via Stochastic Policy Optimization
by: Chen, Yifei, et al.
Published: (2026)
by: Chen, Yifei, et al.
Published: (2026)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
by: Zhang, Jinhao, et al.
Published: (2025)
by: Zhang, Jinhao, et al.
Published: (2025)
Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
by: Wei, Quan, et al.
Published: (2025)
by: Wei, Quan, et al.
Published: (2025)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
by: Liu, Zeyuan, et al.
Published: (2026)
by: Liu, Zeyuan, et al.
Published: (2026)
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
by: Hou, Hongru, et al.
Published: (2026)
by: Hou, Hongru, et al.
Published: (2026)
Personalized Interpolation: Achieving Efficient Conversion Estimation with Flexible Optimization Windows
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Implicit Neural Differential Model for Spatiotemporal Dynamics
by: Akhare, Deepak, et al.
Published: (2025)
by: Akhare, Deepak, et al.
Published: (2025)
Similar Items
-
Wukong: Towards a Scaling Law for Large-Scale Recommendation
by: Zhang, Buyun, et al.
Published: (2024) -
Disaggregated Multi-Tower: Topology-aware Modeling Technique for Efficient Large-Scale Recommendation
by: Luo, Liang, et al.
Published: (2024) -
Proactive Constrained Policy Optimization with Preemptive Penalty
by: Yang, Ning, et al.
Published: (2025) -
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
by: Li, Jiaqian, et al.
Published: (2026) -
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
by: Li, Xiang, et al.
Published: (2026)