Policy Gradient Guidance Enables Test Time Control
Fuente:
arXiv
Guardado en:
| Autores principales: | Qi, Jianing, Tang, Hao, Zhu, Zhigang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
por: Qi, Jianing, et al.
Publicado: (2024)
por: Qi, Jianing, et al.
Publicado: (2024)
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance
por: He, Yufei, et al.
Publicado: (2025)
por: He, Yufei, et al.
Publicado: (2025)
Diffusion Guidance Is a Controllable Policy Improvement Operator
por: Frans, Kevin, et al.
Publicado: (2025)
por: Frans, Kevin, et al.
Publicado: (2025)
Reusing Trajectories in Policy Gradients Enables Fast Convergence
por: Montenegro, Alessandro, et al.
Publicado: (2025)
por: Montenegro, Alessandro, et al.
Publicado: (2025)
A Policy Gradient-Based Sequence-to-Sequence Method for Time Series Prediction
por: Sima, Qi, et al.
Publicado: (2024)
por: Sima, Qi, et al.
Publicado: (2024)
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
por: Li, Hengli, et al.
Publicado: (2025)
por: Li, Hengli, et al.
Publicado: (2025)
Calibrated Test-Time Guidance for Bayesian Inference
por: Geyfman, Daniel, et al.
Publicado: (2026)
por: Geyfman, Daniel, et al.
Publicado: (2026)
Data-Efficient RLVR via Off-Policy Influence Guidance
por: Zhu, Erle, et al.
Publicado: (2025)
por: Zhu, Erle, et al.
Publicado: (2025)
Wasserstein Proximal Policy Gradient
por: Zhu, Zhaoyu, et al.
Publicado: (2026)
por: Zhu, Zhaoyu, et al.
Publicado: (2026)
Learning to Generate Gradients for Test-Time Adaptation via Test-Time Training Layers
por: Deng, Qi, et al.
Publicado: (2024)
por: Deng, Qi, et al.
Publicado: (2024)
Optimisation of Structured Neural Controller Based on Continuous-Time Policy Gradient
por: Cho, Namhoon, et al.
Publicado: (2022)
por: Cho, Namhoon, et al.
Publicado: (2022)
DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QA
por: Yin, Jianing, et al.
Publicado: (2026)
por: Yin, Jianing, et al.
Publicado: (2026)
Gradient Guidance for Diffusion Models: An Optimization Perspective
por: Guo, Yingqing, et al.
Publicado: (2024)
por: Guo, Yingqing, et al.
Publicado: (2024)
Hierarchical Prompt Decision Transformer: Improving Few-Shot Policy Generalization with Global and Adaptive Guidance
por: Wang, Zhe, et al.
Publicado: (2024)
por: Wang, Zhe, et al.
Publicado: (2024)
Doubly Robust Fusion of Many Treatments for Policy Learning
por: Zhu, Ke, et al.
Publicado: (2025)
por: Zhu, Ke, et al.
Publicado: (2025)
Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control
por: Ren, Hao, et al.
Publicado: (2026)
por: Ren, Hao, et al.
Publicado: (2026)
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
por: Kim, Yeongmin, et al.
Publicado: (2026)
por: Kim, Yeongmin, et al.
Publicado: (2026)
Model-Free $δ$-Policy Iteration Based on Damped Newton Method for Nonlinear Continuous-Time H$\infty$ Tracking Control
por: Wang, Qi
Publicado: (2024)
por: Wang, Qi
Publicado: (2024)
ZOTTA: Test-Time Adaptation with Gradient-Free Zeroth-Order Optimization
por: Zhang, Ronghao, et al.
Publicado: (2026)
por: Zhang, Ronghao, et al.
Publicado: (2026)
Group Policy Gradient
por: Chen, Junhua, et al.
Publicado: (2025)
por: Chen, Junhua, et al.
Publicado: (2025)
Robust Control with Gradient Uncertainty
por: Qi, Qian
Publicado: (2025)
por: Qi, Qian
Publicado: (2025)
Algorithm-Relative Trajectory Valuation in Policy Gradient Control
por: Li, Shihao, et al.
Publicado: (2025)
por: Li, Shihao, et al.
Publicado: (2025)
Controllable Continual Test-Time Adaptation
por: Shi, Ziqi, et al.
Publicado: (2024)
por: Shi, Ziqi, et al.
Publicado: (2024)
Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning
por: Zhu, Zhaoyu, et al.
Publicado: (2026)
por: Zhu, Zhaoyu, et al.
Publicado: (2026)
Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-Statistic
por: Zhou, Hongyi, et al.
Publicado: (2026)
por: Zhou, Hongyi, et al.
Publicado: (2026)
When Test-Time Guidance Is Enough: Fast Image and Video Editing with Diffusion Guidance
por: Ghorbel, Ahmed, et al.
Publicado: (2026)
por: Ghorbel, Ahmed, et al.
Publicado: (2026)
TRAM: Test-Time Risk Adaptation with Mixture of Agents
por: Chehade, Mohamad Fares El Hajj, et al.
Publicado: (2024)
por: Chehade, Mohamad Fares El Hajj, et al.
Publicado: (2024)
Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with $k$-step Policy Gradients
por: DeWeese, Alex, et al.
Publicado: (2026)
por: DeWeese, Alex, et al.
Publicado: (2026)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
por: Corrado, Nicholas E., et al.
Publicado: (2023)
por: Corrado, Nicholas E., et al.
Publicado: (2023)
Aligning Flow Map Policies with Optimal Q-Guidance
por: Ziakas, Christos, et al.
Publicado: (2026)
por: Ziakas, Christos, et al.
Publicado: (2026)
Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching
por: Khamis, Alaa, et al.
Publicado: (2026)
por: Khamis, Alaa, et al.
Publicado: (2026)
Differentially Private Policy Gradient
por: Rio, Alexandre, et al.
Publicado: (2025)
por: Rio, Alexandre, et al.
Publicado: (2025)
Functional Natural Policy Gradients
por: Bibaut, Aurelien, et al.
Publicado: (2026)
por: Bibaut, Aurelien, et al.
Publicado: (2026)
Efficient Soft Actor-Critic with LLM-Based Action-Level Guidance for Continuous Control
por: Ma, Hao, et al.
Publicado: (2026)
por: Ma, Hao, et al.
Publicado: (2026)
Learning Optimal Deterministic Policies with Stochastic Policy Gradients
por: Montenegro, Alessandro, et al.
Publicado: (2024)
por: Montenegro, Alessandro, et al.
Publicado: (2024)
When Do Off-Policy and On-Policy Policy Gradient Methods Align?
por: Mambelli, Davide, et al.
Publicado: (2024)
por: Mambelli, Davide, et al.
Publicado: (2024)
Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
por: Jin, Luozhijie, et al.
Publicado: (2025)
por: Jin, Luozhijie, et al.
Publicado: (2025)
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
por: Zhang, Zuyuan, et al.
Publicado: (2026)
por: Zhang, Zuyuan, et al.
Publicado: (2026)
$\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space
por: Wang, Peihao, et al.
Publicado: (2026)
por: Wang, Peihao, et al.
Publicado: (2026)
Residual Policy Gradient: A Reward View of KL-regularized Objective
por: Wang, Pengcheng, et al.
Publicado: (2025)
por: Wang, Pengcheng, et al.
Publicado: (2025)
Ejemplares similares
-
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
por: Qi, Jianing, et al.
Publicado: (2024) -
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance
por: He, Yufei, et al.
Publicado: (2025) -
Diffusion Guidance Is a Controllable Policy Improvement Operator
por: Frans, Kevin, et al.
Publicado: (2025) -
Reusing Trajectories in Policy Gradients Enables Fast Convergence
por: Montenegro, Alessandro, et al.
Publicado: (2025) -
A Policy Gradient-Based Sequence-to-Sequence Method for Time Series Prediction
por: Sima, Qi, et al.
Publicado: (2024)