Saved in:
| Main Author: | Kaddour, Jean |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.06159 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synthetic Data Generation in Low-Resource Settings via Fine-Tuning of Large Language Models
by: Kaddour, Jean, et al.
Published: (2023)
by: Kaddour, Jean, et al.
Published: (2023)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
Causal Machine Learning: A Survey and Open Problems
by: Kaddour, Jean, et al.
Published: (2022)
by: Kaddour, Jean, et al.
Published: (2022)
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026)
by: Kaddour, Jean, et al.
Published: (2026)
REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
by: Stojanovski, Zafir, et al.
Published: (2025)
by: Stojanovski, Zafir, et al.
Published: (2025)
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
Diffusion Policy Policy Optimization
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
by: Shankar, Kaaustaaub, et al.
Published: (2025)
by: Shankar, Kaaustaaub, et al.
Published: (2025)
Simple Policy Optimization
by: Xie, Zhengpeng, et al.
Published: (2024)
by: Xie, Zhengpeng, et al.
Published: (2024)
Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL
by: Sandhu, Dillon, et al.
Published: (2026)
by: Sandhu, Dillon, et al.
Published: (2026)
Think Outside the Policy: In-Context Steered Policy Optimization
by: Huang, Hsiu-Yuan, et al.
Published: (2025)
by: Huang, Hsiu-Yuan, et al.
Published: (2025)
Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
by: Mroueh, Youssef, et al.
Published: (2025)
by: Mroueh, Youssef, et al.
Published: (2025)
Near-Future Policy Optimization
by: Qin, Chuanyu, et al.
Published: (2026)
by: Qin, Chuanyu, et al.
Published: (2026)
Dual Approximation Policy Optimization
by: Xiong, Zhihan, et al.
Published: (2024)
by: Xiong, Zhihan, et al.
Published: (2024)
Reinforcement Learning with a Focus on Adjusting Policies to Reach Targets
by: Tsuboya, Akane, et al.
Published: (2024)
by: Tsuboya, Akane, et al.
Published: (2024)
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
Orthogonalized Policy Optimization:Policy Optimization as Orthogonal Projection in Hilbert Space
by: Zixian, Wang
Published: (2026)
by: Zixian, Wang
Published: (2026)
Diffusion Policy through Conditional Proximal Policy Optimization
by: Liu, Ben, et al.
Published: (2026)
by: Liu, Ben, et al.
Published: (2026)
RePO: Bridging On-Policy Learning and Off-Policy Knowledge through Rephrasing Policy Optimization
by: Xia, Linxuan, et al.
Published: (2026)
by: Xia, Linxuan, et al.
Published: (2026)
Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation
by: Zhang, Xiaoying, et al.
Published: (2024)
by: Zhang, Xiaoying, et al.
Published: (2024)
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
by: Kiyohara, Haruka, et al.
Published: (2024)
by: Kiyohara, Haruka, et al.
Published: (2024)
Wasserstein Policy Optimization
by: Pfau, David, et al.
Published: (2025)
by: Pfau, David, et al.
Published: (2025)
Reflective Policy Optimization
by: Gan, Yaozhong, et al.
Published: (2024)
by: Gan, Yaozhong, et al.
Published: (2024)
Relative Entropy Pathwise Policy Optimization
by: Voelcker, Claas, et al.
Published: (2025)
by: Voelcker, Claas, et al.
Published: (2025)
Central Path Proximal Policy Optimization
by: Milosevic, Nikola, et al.
Published: (2025)
by: Milosevic, Nikola, et al.
Published: (2025)
Dichotomous Diffusion Policy Optimization
by: Liang, Ruiming, et al.
Published: (2025)
by: Liang, Ruiming, et al.
Published: (2025)
Zeroth-Order Optimization is Secretly Single-Step Policy Optimization
by: Qiu, Junbin, et al.
Published: (2025)
by: Qiu, Junbin, et al.
Published: (2025)
Stable On-Policy Distillation through Adaptive Target Reformulation
by: Jang, Ijun, et al.
Published: (2026)
by: Jang, Ijun, et al.
Published: (2026)
Relative Policy-Transition Optimization for Fast Policy Transfer
by: Xu, Jiawei, et al.
Published: (2022)
by: Xu, Jiawei, et al.
Published: (2022)
Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space
by: Zixian, Wang
Published: (2026)
by: Zixian, Wang
Published: (2026)
GRPOformer: Advancing Hyperparameter Optimization via Group Relative Policy Optimization
by: Guo, Haoxin, et al.
Published: (2025)
by: Guo, Haoxin, et al.
Published: (2025)
Policy Targeting under Network Interference
by: Viviano, Davide
Published: (2019)
by: Viviano, Davide
Published: (2019)
Adaptive Simulation Experiment for LLM Policy Optimization
by: Hu, Mingjie, et al.
Published: (2026)
by: Hu, Mingjie, et al.
Published: (2026)
Optimal Regret for Policy Optimization in Contextual Bandits
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
Actor-Critic Pretraining for Proximal Policy Optimization
by: Kernbach, Andreas, et al.
Published: (2026)
by: Kernbach, Andreas, et al.
Published: (2026)
Robust Policy Optimization to Prevent Catastrophic Forgetting
by: Sabbaghi, Mahdi, et al.
Published: (2026)
by: Sabbaghi, Mahdi, et al.
Published: (2026)
Towards Causal Model-Based Policy Optimization
by: Caron, Alberto, et al.
Published: (2025)
by: Caron, Alberto, et al.
Published: (2025)
Transductive Off-policy Proximal Policy Optimization
by: Gan, Yaozhong, et al.
Published: (2024)
by: Gan, Yaozhong, et al.
Published: (2024)
e-COP : Episodic Constrained Optimization of Policies
by: Agnihotri, Akhil, et al.
Published: (2024)
by: Agnihotri, Akhil, et al.
Published: (2024)
Deep Gaussian Process Proximal Policy Optimization
by: van der Lende, Matthijs, et al.
Published: (2025)
by: van der Lende, Matthijs, et al.
Published: (2025)
Similar Items
-
Synthetic Data Generation in Low-Resource Settings via Fine-Tuning of Large Language Models
by: Kaddour, Jean, et al.
Published: (2023) -
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024) -
Causal Machine Learning: A Survey and Open Problems
by: Kaddour, Jean, et al.
Published: (2022) -
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026) -
REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
by: Stojanovski, Zafir, et al.
Published: (2025)