Policy Gradient Guidance Enables Test Time Control
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Jianing, Tang, Hao, Zhu, Zhigang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
by: Qi, Jianing, et al.
Published: (2024)
by: Qi, Jianing, et al.
Published: (2024)
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance
by: He, Yufei, et al.
Published: (2025)
by: He, Yufei, et al.
Published: (2025)
Diffusion Guidance Is a Controllable Policy Improvement Operator
by: Frans, Kevin, et al.
Published: (2025)
by: Frans, Kevin, et al.
Published: (2025)
Reusing Trajectories in Policy Gradients Enables Fast Convergence
by: Montenegro, Alessandro, et al.
Published: (2025)
by: Montenegro, Alessandro, et al.
Published: (2025)
A Policy Gradient-Based Sequence-to-Sequence Method for Time Series Prediction
by: Sima, Qi, et al.
Published: (2024)
by: Sima, Qi, et al.
Published: (2024)
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
by: Li, Hengli, et al.
Published: (2025)
by: Li, Hengli, et al.
Published: (2025)
Calibrated Test-Time Guidance for Bayesian Inference
by: Geyfman, Daniel, et al.
Published: (2026)
by: Geyfman, Daniel, et al.
Published: (2026)
Data-Efficient RLVR via Off-Policy Influence Guidance
by: Zhu, Erle, et al.
Published: (2025)
by: Zhu, Erle, et al.
Published: (2025)
Wasserstein Proximal Policy Gradient
by: Zhu, Zhaoyu, et al.
Published: (2026)
by: Zhu, Zhaoyu, et al.
Published: (2026)
Learning to Generate Gradients for Test-Time Adaptation via Test-Time Training Layers
by: Deng, Qi, et al.
Published: (2024)
by: Deng, Qi, et al.
Published: (2024)
Optimisation of Structured Neural Controller Based on Continuous-Time Policy Gradient
by: Cho, Namhoon, et al.
Published: (2022)
by: Cho, Namhoon, et al.
Published: (2022)
DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QA
by: Yin, Jianing, et al.
Published: (2026)
by: Yin, Jianing, et al.
Published: (2026)
Gradient Guidance for Diffusion Models: An Optimization Perspective
by: Guo, Yingqing, et al.
Published: (2024)
by: Guo, Yingqing, et al.
Published: (2024)
Hierarchical Prompt Decision Transformer: Improving Few-Shot Policy Generalization with Global and Adaptive Guidance
by: Wang, Zhe, et al.
Published: (2024)
by: Wang, Zhe, et al.
Published: (2024)
Doubly Robust Fusion of Many Treatments for Policy Learning
by: Zhu, Ke, et al.
Published: (2025)
by: Zhu, Ke, et al.
Published: (2025)
Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control
by: Ren, Hao, et al.
Published: (2026)
by: Ren, Hao, et al.
Published: (2026)
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
by: Kim, Yeongmin, et al.
Published: (2026)
by: Kim, Yeongmin, et al.
Published: (2026)
Model-Free $δ$-Policy Iteration Based on Damped Newton Method for Nonlinear Continuous-Time H$\infty$ Tracking Control
by: Wang, Qi
Published: (2024)
by: Wang, Qi
Published: (2024)
ZOTTA: Test-Time Adaptation with Gradient-Free Zeroth-Order Optimization
by: Zhang, Ronghao, et al.
Published: (2026)
by: Zhang, Ronghao, et al.
Published: (2026)
Group Policy Gradient
by: Chen, Junhua, et al.
Published: (2025)
by: Chen, Junhua, et al.
Published: (2025)
Robust Control with Gradient Uncertainty
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Algorithm-Relative Trajectory Valuation in Policy Gradient Control
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
Controllable Continual Test-Time Adaptation
by: Shi, Ziqi, et al.
Published: (2024)
by: Shi, Ziqi, et al.
Published: (2024)
Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning
by: Zhu, Zhaoyu, et al.
Published: (2026)
by: Zhu, Zhaoyu, et al.
Published: (2026)
Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-Statistic
by: Zhou, Hongyi, et al.
Published: (2026)
by: Zhou, Hongyi, et al.
Published: (2026)
When Test-Time Guidance Is Enough: Fast Image and Video Editing with Diffusion Guidance
by: Ghorbel, Ahmed, et al.
Published: (2026)
by: Ghorbel, Ahmed, et al.
Published: (2026)
TRAM: Test-Time Risk Adaptation with Mixture of Agents
by: Chehade, Mohamad Fares El Hajj, et al.
Published: (2024)
by: Chehade, Mohamad Fares El Hajj, et al.
Published: (2024)
Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with $k$-step Policy Gradients
by: DeWeese, Alex, et al.
Published: (2026)
by: DeWeese, Alex, et al.
Published: (2026)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
by: Corrado, Nicholas E., et al.
Published: (2023)
by: Corrado, Nicholas E., et al.
Published: (2023)
Aligning Flow Map Policies with Optimal Q-Guidance
by: Ziakas, Christos, et al.
Published: (2026)
by: Ziakas, Christos, et al.
Published: (2026)
Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching
by: Khamis, Alaa, et al.
Published: (2026)
by: Khamis, Alaa, et al.
Published: (2026)
Differentially Private Policy Gradient
by: Rio, Alexandre, et al.
Published: (2025)
by: Rio, Alexandre, et al.
Published: (2025)
Functional Natural Policy Gradients
by: Bibaut, Aurelien, et al.
Published: (2026)
by: Bibaut, Aurelien, et al.
Published: (2026)
Efficient Soft Actor-Critic with LLM-Based Action-Level Guidance for Continuous Control
by: Ma, Hao, et al.
Published: (2026)
by: Ma, Hao, et al.
Published: (2026)
Learning Optimal Deterministic Policies with Stochastic Policy Gradients
by: Montenegro, Alessandro, et al.
Published: (2024)
by: Montenegro, Alessandro, et al.
Published: (2024)
When Do Off-Policy and On-Policy Policy Gradient Methods Align?
by: Mambelli, Davide, et al.
Published: (2024)
by: Mambelli, Davide, et al.
Published: (2024)
Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
by: Jin, Luozhijie, et al.
Published: (2025)
by: Jin, Luozhijie, et al.
Published: (2025)
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
$\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space
by: Wang, Peihao, et al.
Published: (2026)
by: Wang, Peihao, et al.
Published: (2026)
Residual Policy Gradient: A Reward View of KL-regularized Objective
by: Wang, Pengcheng, et al.
Published: (2025)
by: Wang, Pengcheng, et al.
Published: (2025)
Similar Items
-
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
by: Qi, Jianing, et al.
Published: (2024) -
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance
by: He, Yufei, et al.
Published: (2025) -
Diffusion Guidance Is a Controllable Policy Improvement Operator
by: Frans, Kevin, et al.
Published: (2025) -
Reusing Trajectories in Policy Gradients Enables Fast Convergence
by: Montenegro, Alessandro, et al.
Published: (2025) -
A Policy Gradient-Based Sequence-to-Sequence Method for Time Series Prediction
by: Sima, Qi, et al.
Published: (2024)