Policy Gradient with Kernel Quadrature
Fuente:
arXiv
Saved in:
| Main Authors: | Hayakawa, Satoshi, Morimura, Tetsuro |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes
by: Morimura, Tetsuro, et al.
Published: (2022)
by: Morimura, Tetsuro, et al.
Published: (2022)
Interaction Locality in Hierarchical Recursive Reasoning
by: Miyanishi, Yosuke, et al.
Published: (2026)
by: Miyanishi, Yosuke, et al.
Published: (2026)
Convex-Geometric Error Bounds for Positive-Weight Kernel Quadrature
by: Hayakawa, Satoshi
Published: (2026)
by: Hayakawa, Satoshi
Published: (2026)
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024)
by: Morimura, Tetsuro, et al.
Published: (2024)
Conformal Prediction as Bayesian Quadrature
by: Snell, Jake C., et al.
Published: (2025)
by: Snell, Jake C., et al.
Published: (2025)
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025)
by: Ståhlberg, Simon, et al.
Published: (2025)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Policy Newton Algorithm in Reproducing Kernel Hilbert Space
by: Zhang, Yixian, et al.
Published: (2025)
by: Zhang, Yixian, et al.
Published: (2025)
Geometrically Inspired Kernel Machines for Collaborative Learning Beyond Gradient Descent
by: Kumar, Mohit, et al.
Published: (2024)
by: Kumar, Mohit, et al.
Published: (2024)
Gradient Extrapolation-Based Policy Optimization
by: Swapnil, Ismam Nur, et al.
Published: (2026)
by: Swapnil, Ismam Nur, et al.
Published: (2026)
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
Measures of Variability for Risk-averse Policy Gradient
by: Luo, Yudong, et al.
Published: (2025)
by: Luo, Yudong, et al.
Published: (2025)
Soft Deterministic Policy Gradient with Gaussian Smoothing
by: Na, Hyunjun, et al.
Published: (2026)
by: Na, Hyunjun, et al.
Published: (2026)
Towards Provable Log Density Policy Gradient
by: Katdare, Pulkit, et al.
Published: (2024)
by: Katdare, Pulkit, et al.
Published: (2024)
Policy Gradient for Robust Markov Decision Processes
by: Wang, Qiuhao, et al.
Published: (2024)
by: Wang, Qiuhao, et al.
Published: (2024)
Sequential Policy Gradient for Adaptive Hyperparameter Optimization
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
by: Onoda, Ku, et al.
Published: (2026)
by: Onoda, Ku, et al.
Published: (2026)
Policy Gradient Methods in the Presence of Symmetries and State Abstractions
by: Panangaden, Prakash, et al.
Published: (2023)
by: Panangaden, Prakash, et al.
Published: (2023)
Vertical Symbolic Regression via Deep Policy Gradient
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024)
by: Pan, Deng, et al.
Published: (2024)
Policy Gradient Methods for Non-Markovian Reinforcement Learning
by: Kar, Avik, et al.
Published: (2026)
by: Kar, Avik, et al.
Published: (2026)
GRADE: Replacing Policy Gradients with Backpropagation for LLM Alignment
by: Nel, Lukas Abrie
Published: (2025)
by: Nel, Lukas Abrie
Published: (2025)
Matrix Low-Rank Approximation For Policy Gradient Methods
by: Rozada, Sergio, et al.
Published: (2024)
by: Rozada, Sergio, et al.
Published: (2024)
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
by: Lepel, Olivier, et al.
Published: (2024)
by: Lepel, Olivier, et al.
Published: (2024)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
by: Meng, Wenjia, et al.
Published: (2024)
by: Meng, Wenjia, et al.
Published: (2024)
Delightful Policy Gradient
by: Osband, Ian
Published: (2026)
by: Osband, Ian
Published: (2026)
Mollification Effects of Policy Gradient Methods
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement Learning
by: Batra, Sumeet, et al.
Published: (2023)
by: Batra, Sumeet, et al.
Published: (2023)
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
by: Mousavi-Hosseini, Alireza, et al.
Published: (2026)
by: Mousavi-Hosseini, Alireza, et al.
Published: (2026)
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
by: Nishimori, Soichiro, et al.
Published: (2026)
by: Nishimori, Soichiro, et al.
Published: (2026)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
by: Reddi, Aryaman, et al.
Published: (2025)
by: Reddi, Aryaman, et al.
Published: (2025)
Policy Gradient with Adaptive Entropy Annealing for Continual Fine-Tuning
by: Zhang, Yaqian, et al.
Published: (2026)
by: Zhang, Yaqian, et al.
Published: (2026)
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic
by: Lee, Jeong Woon, et al.
Published: (2026)
by: Lee, Jeong Woon, et al.
Published: (2026)
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2025)
by: Galesloot, Maris F. L., et al.
Published: (2025)
FlowPG: Action-constrained Policy Gradient with Normalizing Flows
by: Brahmanage, Janaka Chathuranga, et al.
Published: (2024)
by: Brahmanage, Janaka Chathuranga, et al.
Published: (2024)
Do Transformer World Models Give Better Policy Gradients?
by: Ma, Michel, et al.
Published: (2024)
by: Ma, Michel, et al.
Published: (2024)
Delightful Distributed Policy Gradient
by: Osband, Ian
Published: (2026)
by: Osband, Ian
Published: (2026)
Similar Items
-
Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes
by: Morimura, Tetsuro, et al.
Published: (2022) -
Interaction Locality in Hierarchical Recursive Reasoning
by: Miyanishi, Yosuke, et al.
Published: (2026) -
Convex-Geometric Error Bounds for Positive-Weight Kernel Quadrature
by: Hayakawa, Satoshi
Published: (2026) -
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024) -
Conformal Prediction as Bayesian Quadrature
by: Snell, Jake C., et al.
Published: (2025)