Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Zhongzhu, Yang, Yibo, Chen, Ziyan, Bie, Fengxiang, Xia, Haojun, Wu, Xiaoxia, Wu, Robert, Athiwaratkun, Ben, Ghanem, Bernard, Song, Shuaiwen Leon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning
by: Yang, Yibo, et al.
Published: (2024)
by: Yang, Yibo, et al.
Published: (2024)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
by: Xia, Haojun, et al.
Published: (2025)
by: Xia, Haojun, et al.
Published: (2025)
When RL Meets Adaptive Speculative Training: A Unified Training-Serving System
by: Wang, Junxiong, et al.
Published: (2026)
by: Wang, Junxiong, et al.
Published: (2026)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
by: Jia, Jinda, et al.
Published: (2026)
by: Jia, Jinda, et al.
Published: (2026)
Data Diversification Methods In Alignment Enhance Math Performance In LLMs
by: Dokmeci, Berkan, et al.
Published: (2025)
by: Dokmeci, Berkan, et al.
Published: (2025)
Introspective Diffusion Language Models
by: Yu, Yifan, et al.
Published: (2026)
by: Yu, Yifan, et al.
Published: (2026)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
by: Xia, Haojun, et al.
Published: (2024)
by: Xia, Haojun, et al.
Published: (2024)
Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping
by: Zhang, Muru, et al.
Published: (2025)
by: Zhang, Muru, et al.
Published: (2025)
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
by: Thapa, Rahul, et al.
Published: (2024)
by: Thapa, Rahul, et al.
Published: (2024)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
by: Meng, Wenjia, et al.
Published: (2024)
by: Meng, Wenjia, et al.
Published: (2024)
Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
by: Thapa, Rahul, et al.
Published: (2025)
by: Thapa, Rahul, et al.
Published: (2025)
English‐Only Policies and Allegations of Racism in Nursing: Safety, Culture and Respect Prevail
by: Sharon Brownie, et al.
Published: (2025)
by: Sharon Brownie, et al.
Published: (2025)
Learning Optimal Deterministic Policies with Stochastic Policy Gradients
by: Montenegro, Alessandro, et al.
Published: (2024)
by: Montenegro, Alessandro, et al.
Published: (2024)
Purifying Task Vectors in Knowledge-Aware Subspace for Model Merging
by: An, Bang, et al.
Published: (2025)
by: An, Bang, et al.
Published: (2025)
Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping
by: Snyder, David, et al.
Published: (2025)
by: Snyder, David, et al.
Published: (2025)
Disentangling Reasoning and Knowledge in Medical Large Language Models
by: Thapa, Rahul, et al.
Published: (2025)
by: Thapa, Rahul, et al.
Published: (2025)
GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation
by: Alhuwaider, Shyma, et al.
Published: (2026)
by: Alhuwaider, Shyma, et al.
Published: (2026)
Towards Interpretable Deep Local Learning with Successive Gradient Reconciliation
by: Yang, Yibo, et al.
Published: (2024)
by: Yang, Yibo, et al.
Published: (2024)
Video Self-Stitching Graph Network for Temporal Action Localization
by: Zhao, Chen, et al.
Published: (2020)
by: Zhao, Chen, et al.
Published: (2020)
Geometry-aware Policy Imitation
by: Li, Yiming, et al.
Published: (2025)
by: Li, Yiming, et al.
Published: (2025)
Policy Gradients for Optimal Parallel Tempering MCMC
by: Zhao, Daniel, et al.
Published: (2024)
by: Zhao, Daniel, et al.
Published: (2024)
HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning
by: Zhang, Jiyao, et al.
Published: (2026)
by: Zhang, Jiyao, et al.
Published: (2026)
Policy Gradient Converges to the Globally Optimal Policy for Nearly Linear-Quadratic Regulators
by: Han, Yinbin, et al.
Published: (2023)
by: Han, Yinbin, et al.
Published: (2023)
GFlowNet Training by Policy Gradients
by: Niu, Puhua, et al.
Published: (2024)
by: Niu, Puhua, et al.
Published: (2024)
CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation
by: Songwei, Wu, et al.
Published: (2026)
by: Songwei, Wu, et al.
Published: (2026)
Human-assisted Robotic Policy Refinement via Action Preference Optimization
by: Xia, Wenke, et al.
Published: (2025)
by: Xia, Wenke, et al.
Published: (2025)
Globally Stable Neural Imitation Policies
by: Abyaneh, Amin, et al.
Published: (2024)
by: Abyaneh, Amin, et al.
Published: (2024)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
by: Xu, Huimin, et al.
Published: (2026)
by: Xu, Huimin, et al.
Published: (2026)
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
Overcoming Valid Action Suppression in Unmasked Policy Gradient Algorithms
by: Zabounidis, Renos, et al.
Published: (2026)
by: Zabounidis, Renos, et al.
Published: (2026)
FlowPG: Action-constrained Policy Gradient with Normalizing Flows
by: Brahmanage, Janaka Chathuranga, et al.
Published: (2024)
by: Brahmanage, Janaka Chathuranga, et al.
Published: (2024)
Zero Collapse: A Failure Mode of Policy Gradient Methods in Discontinuous Reward Environments
by: Kumar, Nishant, et al.
Published: (2026)
by: Kumar, Nishant, et al.
Published: (2026)
AED: Adaptable Error Detection for Few-shot Imitation Policy
by: Yeh, Jia-Fong, et al.
Published: (2024)
by: Yeh, Jia-Fong, et al.
Published: (2024)
Uncertainty-Aware Deployment of Pre-trained Language-Conditioned Imitation Learning Policies
by: Wu, Bo, et al.
Published: (2024)
by: Wu, Bo, et al.
Published: (2024)
Where-to-Learn: Analytical Policy Gradient Directed Exploration for On-Policy Robotic Reinforcement Learning
by: Chang, Leixin, et al.
Published: (2026)
by: Chang, Leixin, et al.
Published: (2026)
Similar Items
-
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
by: Zhou, Zhongzhu, et al.
Published: (2026) -
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
by: Zhou, Zhongzhu, et al.
Published: (2026) -
CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning
by: Yang, Yibo, et al.
Published: (2024) -
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
by: Zhang, Zhenyu, et al.
Published: (2025) -
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
by: Xia, Haojun, et al.
Published: (2025)