Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chenliang, Elmahdy, Adel, Boyd, Alex, Wang, Zhongruo, Zeng, Siliang, Garcia, Alfredo, Bhatia, Parminder, Kass-Hout, Taha, Xiao, Cao, Hong, Mingyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hyper Hawkes Processes: Interpretable Models of Marked Temporal Point Processes
by: Boyd, Alex, et al.
Published: (2025)
by: Boyd, Alex, et al.
Published: (2025)
Deep Continuous-Time State-Space Models for Marked Event Sequences
by: Chang, Yuxin, et al.
Published: (2024)
by: Chang, Yuxin, et al.
Published: (2024)
MammoDINO: Anatomically Aware Self-Supervision for Mammographic Images
by: Zhou, Sicheng, et al.
Published: (2025)
by: Zhou, Sicheng, et al.
Published: (2025)
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning
by: Chang, Aofei, et al.
Published: (2025)
by: Chang, Aofei, et al.
Published: (2025)
MedHEval: Benchmarking Hallucinations and Mitigation Strategies in Medical Large Vision-Language Models
by: Chang, Aofei, et al.
Published: (2025)
by: Chang, Aofei, et al.
Published: (2025)
Dynamic Uncertainty Ranking: Enhancing Retrieval-Augmented In-Context Learning for Long-Tail Knowledge in LLMs
by: Yu, Shuyang, et al.
Published: (2024)
by: Yu, Shuyang, et al.
Published: (2024)
Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector
by: Yang, Haoyan, et al.
Published: (2025)
by: Yang, Haoyan, et al.
Published: (2025)
Segment as You Wish -- Free-Form Language-Based Segmentation for Medical Images
by: Da, Longchao, et al.
Published: (2024)
by: Da, Longchao, et al.
Published: (2024)
Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image Segmentation
by: Konwer, Aishik, et al.
Published: (2025)
by: Konwer, Aishik, et al.
Published: (2025)
Bi-level Contrastive Learning for Knowledge-Enhanced Molecule Representations
by: Jiang, Pengcheng, et al.
Published: (2023)
by: Jiang, Pengcheng, et al.
Published: (2023)
Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval
by: Jiang, Pengcheng, et al.
Published: (2024)
by: Jiang, Pengcheng, et al.
Published: (2024)
When Demonstrations Meet Generative World Models: A Maximum Likelihood Framework for Offline Inverse Reinforcement Learning
by: Zeng, Siliang, et al.
Published: (2023)
by: Zeng, Siliang, et al.
Published: (2023)
Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
by: Wei, Quan, et al.
Published: (2025)
by: Wei, Quan, et al.
Published: (2025)
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
by: Zhang, Xinnan, et al.
Published: (2025)
by: Zhang, Xinnan, et al.
Published: (2025)
Understanding Inverse Reinforcement Learning under Overparameterization: Non-Asymptotic Analysis and Global Optimality
by: Zhang, Ruijia, et al.
Published: (2025)
by: Zhang, Ruijia, et al.
Published: (2025)
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment
by: Li, Chenliang, et al.
Published: (2024)
by: Li, Chenliang, et al.
Published: (2024)
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
by: Li, Jiaxiang, et al.
Published: (2024)
by: Li, Jiaxiang, et al.
Published: (2024)
Structural Estimation of Markov Decision Processes in High-Dimensional State Space with Finite-Time Guarantees
by: Zeng, Siliang, et al.
Published: (2022)
by: Zeng, Siliang, et al.
Published: (2022)
A Bayesian Approach to Robust Inverse Reinforcement Learning
by: Wei, Ran, et al.
Published: (2023)
by: Wei, Ran, et al.
Published: (2023)
Synergistic Approach for Simultaneous Optimization of Monolingual, Cross-lingual, and Multilingual Information Retrieval
by: Elmahdy, Adel, et al.
Published: (2024)
by: Elmahdy, Adel, et al.
Published: (2024)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
by: Yang, Zhijian, et al.
Published: (2025)
by: Yang, Zhijian, et al.
Published: (2025)
A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping
by: Chen, Dingwei, et al.
Published: (2026)
by: Chen, Dingwei, et al.
Published: (2026)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
by: Zhou, Hongyi, et al.
Published: (2025)
by: Zhou, Hongyi, et al.
Published: (2025)
Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities
by: Xie, Ziwen, et al.
Published: (2026)
by: Xie, Ziwen, et al.
Published: (2026)
Differentially Private SGD Without Clipping Bias: An Error-Feedback Approach
by: Zhang, Xinwei, et al.
Published: (2023)
by: Zhang, Xinwei, et al.
Published: (2023)
RHRSegNet: Relighting High-Resolution Night-Time Semantic Segmentation
by: Elmahdy, Sarah, et al.
Published: (2024)
by: Elmahdy, Sarah, et al.
Published: (2024)
Policy Gradient with Active Importance Sampling
by: Papini, Matteo, et al.
Published: (2024)
by: Papini, Matteo, et al.
Published: (2024)
Posterior Optimization with Clipped Objective for Bridging Efficiency and Stability in Generative Policy Learning
by: Chen, Yuhui, et al.
Published: (2026)
by: Chen, Yuhui, et al.
Published: (2026)
Low Variance Off-policy Evaluation with State-based Importance Sampling
by: Bossens, David M., et al.
Published: (2022)
by: Bossens, David M., et al.
Published: (2022)
From Gradient Clipping to Normalization for Heavy Tailed SGD
by: Hübler, Florian, et al.
Published: (2024)
by: Hübler, Florian, et al.
Published: (2024)
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
by: Khah, Saleh Vatan, et al.
Published: (2025)
by: Khah, Saleh Vatan, et al.
Published: (2025)
An Investigation of Batch Normalization in Off-Policy Actor-Critic Algorithms
by: Wang, Li, et al.
Published: (2025)
by: Wang, Li, et al.
Published: (2025)
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
ESPO: Entropy Importance Sampling Policy Optimization
by: Sheng, Yuepeng, et al.
Published: (2025)
by: Sheng, Yuepeng, et al.
Published: (2025)
GIPO: Gaussian Importance Sampling Policy Optimization
by: Lu, Chengxuan, et al.
Published: (2026)
by: Lu, Chengxuan, et al.
Published: (2026)
Off Policy Lyapunov Stability in Reinforcement Learning
by: Gill, Sarvan, et al.
Published: (2025)
by: Gill, Sarvan, et al.
Published: (2025)
DCPO: Dynamic Clipping Policy Optimization
by: Yang, Shihui, et al.
Published: (2025)
by: Yang, Shihui, et al.
Published: (2025)
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
“Trees give life. Police take it”: Building and Fighting for Abolitionist Life‐Worlds, from the Weelaunee Forest to Georgia's Jails
by: Hannah Kass
Published: (2025)
by: Hannah Kass
Published: (2025)
Similar Items
-
Hyper Hawkes Processes: Interpretable Models of Marked Temporal Point Processes
by: Boyd, Alex, et al.
Published: (2025) -
Deep Continuous-Time State-Space Models for Marked Event Sequences
by: Chang, Yuxin, et al.
Published: (2024) -
MammoDINO: Anatomically Aware Self-Supervision for Mammographic Images
by: Zhou, Sicheng, et al.
Published: (2025) -
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning
by: Chang, Aofei, et al.
Published: (2025) -
MedHEval: Benchmarking Hallucinations and Mitigation Strategies in Medical Large Vision-Language Models
by: Chang, Aofei, et al.
Published: (2025)