Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
Fuente:
arXiv
Saved in:
| Main Authors: | Guan, Zhong, Guo, Yongjian, Sun, Haoran, Huang, Wen, Di, Shuai, Wu, Likang, Wu, Xiong Jun, Zhao, Hongke |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training
by: Sun, Haoran, et al.
Published: (2026)
by: Sun, Haoran, et al.
Published: (2026)
Enhancing Collaborative Semantics of Language Model-Driven Recommendations via Graph-Aware Learning
by: Guan, Zhong, et al.
Published: (2024)
by: Guan, Zhong, et al.
Published: (2024)
Recall-Extend Dynamics: Enhancing Small Language Models through Controlled Exploration and Refined Offline Integration
by: Guan, Zhong, et al.
Published: (2025)
by: Guan, Zhong, et al.
Published: (2025)
Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data
by: Guan, Zhong, et al.
Published: (2025)
by: Guan, Zhong, et al.
Published: (2025)
LangTopo: Aligning Language Descriptions of Graphs with Tokenized Topological Modeling
by: Guan, Zhong, et al.
Published: (2024)
by: Guan, Zhong, et al.
Published: (2024)
Hierarchical Semantic RL: Tackling the Problem of Dynamic Action Space for RL-based Recommendations
by: Wang, Minmao, et al.
Published: (2025)
by: Wang, Minmao, et al.
Published: (2025)
Reinventing Clinical Dialogue: Agentic Paradigms for LLM Enabled Healthcare Communication
by: Zhi, Xiaoquan, et al.
Published: (2025)
by: Zhi, Xiaoquan, et al.
Published: (2025)
Multi-View Empowered Structural Graph Wordification for Language Models
by: Liu, Zipeng, et al.
Published: (2024)
by: Liu, Zipeng, et al.
Published: (2024)
D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models
by: Guo, Yucheng, et al.
Published: (2026)
by: Guo, Yucheng, et al.
Published: (2026)
Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
by: Wang, Jiacheng, et al.
Published: (2026)
by: Wang, Jiacheng, et al.
Published: (2026)
GANPrompt: Enhancing Robustness in LLM-Based Recommendations with GAN-Enhanced Diversity Prompts
by: Li, Xinyu, et al.
Published: (2024)
by: Li, Xinyu, et al.
Published: (2024)
LANE: Logic Alignment of Non-tuning Large Language Models and Online Recommendation Systems for Explainable Reason Generation
by: Zhao, Hongke, et al.
Published: (2024)
by: Zhao, Hongke, et al.
Published: (2024)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
by: Noukhovitch, Michael, et al.
Published: (2024)
by: Noukhovitch, Michael, et al.
Published: (2024)
AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training
by: Guo, Yucheng, et al.
Published: (2026)
by: Guo, Yucheng, et al.
Published: (2026)
NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models
by: Huang, Wen, et al.
Published: (2026)
by: Huang, Wen, et al.
Published: (2026)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
by: Ye, Chenlu, et al.
Published: (2026)
by: Ye, Chenlu, et al.
Published: (2026)
Differentially Private Distributed Mismatch Tracking Algorithm for Constraint-Coupled Resource Allocation Problems
by: Wu, Wenwen, et al.
Published: (2022)
by: Wu, Wenwen, et al.
Published: (2022)
Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
by: Gao, Jiaxuan, et al.
Published: (2025)
by: Gao, Jiaxuan, et al.
Published: (2025)
Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training
by: Gao, Jiaxuan, et al.
Published: (2026)
by: Gao, Jiaxuan, et al.
Published: (2026)
Robust Contrastive Graph Clustering with Adaptive Local-Global Integration
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
LLMs Can Learn to Reason Via Off-Policy RL
by: Ritter, Daniel, et al.
Published: (2026)
by: Ritter, Daniel, et al.
Published: (2026)
Align and Filter: Improving Performance in Asynchronous On-Policy RL
by: Honari, Homayoun, et al.
Published: (2026)
by: Honari, Homayoun, et al.
Published: (2026)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
Logit Distillation on Manifolds: Mapping by Learning
by: Yang, Yiru, et al.
Published: (2026)
by: Yang, Yiru, et al.
Published: (2026)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025)
by: Ankile, Lars, et al.
Published: (2025)
VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation
by: Sun, Shikun, et al.
Published: (2026)
by: Sun, Shikun, et al.
Published: (2026)
Off-Policy Evaluation for Recommendations with Missing-Not-At-Random Rewards
by: Takahashi, Tatsuki, et al.
Published: (2025)
by: Takahashi, Tatsuki, et al.
Published: (2025)
Off-Policy Evaluation Under Nonignorable Missing Data
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
by: Wang, Taiyi, et al.
Published: (2024)
by: Wang, Taiyi, et al.
Published: (2024)
A Unifying View of Linear Function Approximation in Off-Policy RL Through Matrix Splitting and Preconditioning
by: Wu, Zechen, et al.
Published: (2025)
by: Wu, Zechen, et al.
Published: (2025)
Laminar: A Scalable Asynchronous RL Post-Training Framework
by: Sheng, Guangming, et al.
Published: (2025)
by: Sheng, Guangming, et al.
Published: (2025)
ATP-Dependent Mismatch Recognition in DNA Replication Mismatch Repair
by: Zhang, Nianqin, et al.
Published: (2022)
by: Zhang, Nianqin, et al.
Published: (2022)
Policy Learning for Off-Dynamics RL with Deficient Support
by: Van, Linh Le Pham, et al.
Published: (2024)
by: Van, Linh Le Pham, et al.
Published: (2024)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
Diffmv: A Unified Diffusion Framework for Healthcare Predictions with Random Missing Views and View Laziness
by: Zhao, Chuang, et al.
Published: (2025)
by: Zhao, Chuang, et al.
Published: (2025)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
An Efficient Continuous Control Perspective for Reinforcement-Learning-based Sequential Recommendation
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
AgenticCache: Cache-Driven Asynchronous Planning for Embodied AI Agents
by: Kim, Hojoon, et al.
Published: (2026)
by: Kim, Hojoon, et al.
Published: (2026)
Heddle: A Distributed Orchestration System for Agentic RL Rollout
by: Zhang, Zili, et al.
Published: (2026)
by: Zhang, Zili, et al.
Published: (2026)
Similar Items
-
RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training
by: Sun, Haoran, et al.
Published: (2026) -
Enhancing Collaborative Semantics of Language Model-Driven Recommendations via Graph-Aware Learning
by: Guan, Zhong, et al.
Published: (2024) -
Recall-Extend Dynamics: Enhancing Small Language Models through Controlled Exploration and Refined Offline Integration
by: Guan, Zhong, et al.
Published: (2025) -
Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data
by: Guan, Zhong, et al.
Published: (2025) -
LangTopo: Aligning Language Descriptions of Graphs with Tokenized Topological Modeling
by: Guan, Zhong, et al.
Published: (2024)