BNPO: Beta Normalization Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Changyi, Zhang, Mengdi, Cao, Yixin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Knowledge Graph Embedding by Normalizing Flows
von: Xiao, Changyi, et al.
Veröffentlicht: (2024)
von: Xiao, Changyi, et al.
Veröffentlicht: (2024)
Knowledge Graph Completion by Intermediate Variables Regularization
von: Xiao, Changyi, et al.
Veröffentlicht: (2025)
von: Xiao, Changyi, et al.
Veröffentlicht: (2025)
Complex Logical Query Answering by Calibrating Knowledge Graph Completion Models
von: Xiao, Changyi, et al.
Veröffentlicht: (2024)
von: Xiao, Changyi, et al.
Veröffentlicht: (2024)
Reinforcement Learning with Conditional Expectation Reward
von: Xiao, Changyi, et al.
Veröffentlicht: (2026)
von: Xiao, Changyi, et al.
Veröffentlicht: (2026)
FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect Generalization
von: Hong, Shibo, et al.
Veröffentlicht: (2025)
von: Hong, Shibo, et al.
Veröffentlicht: (2025)
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
von: Gao, Chen-Xiao, et al.
Veröffentlicht: (2025)
von: Gao, Chen-Xiao, et al.
Veröffentlicht: (2025)
ARM: Role-Conditioned Neuron Transplantation for Training-Free Generalist LLM Agent Merging
von: Feng, Zhuoka, et al.
Veröffentlicht: (2026)
von: Feng, Zhuoka, et al.
Veröffentlicht: (2026)
ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical Dialogue
von: Cao, Ruike, et al.
Veröffentlicht: (2026)
von: Cao, Ruike, et al.
Veröffentlicht: (2026)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2026)
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2026)
Does Your Optimizer Care How You Normalize? Normalization-Optimizer Coupling in LLM Training
von: Abouzeid, Abdelrahman
Veröffentlicht: (2026)
von: Abouzeid, Abdelrahman
Veröffentlicht: (2026)
P^2O: Joint Policy and Prompt Optimization
von: Lu, Xinyu, et al.
Veröffentlicht: (2026)
von: Lu, Xinyu, et al.
Veröffentlicht: (2026)
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025)
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025)
Towards Interpretable Reinforcement Learning with Constrained Normalizing Flow Policies
von: Rietz, Finn, et al.
Veröffentlicht: (2024)
von: Rietz, Finn, et al.
Veröffentlicht: (2024)
FlowPG: Action-constrained Policy Gradient with Normalizing Flows
von: Brahmanage, Janaka Chathuranga, et al.
Veröffentlicht: (2024)
von: Brahmanage, Janaka Chathuranga, et al.
Veröffentlicht: (2024)
Relative Policy-Transition Optimization for Fast Policy Transfer
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
Divergence-Augmented Policy Optimization
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Continuous-Time Analysis of Adaptive Optimization and Normalization
von: Gould, Rhys, et al.
Veröffentlicht: (2024)
von: Gould, Rhys, et al.
Veröffentlicht: (2024)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
CROP: Conservative Reward for Model-based Offline Policy Optimization
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
Orthogonalized Policy Optimization:Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck
von: Deng, Huilin, et al.
Veröffentlicht: (2026)
von: Deng, Huilin, et al.
Veröffentlicht: (2026)
Latent Bayesian Optimization via Autoregressive Normalizing Flows
von: Lee, Seunghun, et al.
Veröffentlicht: (2025)
von: Lee, Seunghun, et al.
Veröffentlicht: (2025)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
When Maximum Entropy Misleads Policy Optimization
von: Zhang, Ruipeng, et al.
Veröffentlicht: (2025)
von: Zhang, Ruipeng, et al.
Veröffentlicht: (2025)
Wasserstein Policy Optimization
von: Pfau, David, et al.
Veröffentlicht: (2025)
von: Pfau, David, et al.
Veröffentlicht: (2025)
Reflective Policy Optimization
von: Gan, Yaozhong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozhong, et al.
Veröffentlicht: (2024)
AGPO: Adaptive Group Policy Optimization with Dual Statistical Feedback
von: Hu, Miaobo, et al.
Veröffentlicht: (2026)
von: Hu, Miaobo, et al.
Veröffentlicht: (2026)
ESPO: Entropy Importance Sampling Policy Optimization
von: Sheng, Yuepeng, et al.
Veröffentlicht: (2025)
von: Sheng, Yuepeng, et al.
Veröffentlicht: (2025)
Calibration-Aware Policy Optimization for Reasoning LLMs
von: Wang, Ziqi, et al.
Veröffentlicht: (2026)
von: Wang, Ziqi, et al.
Veröffentlicht: (2026)
UCPO: Uncertainty-Aware Policy Optimization
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
GIPO: Gaussian Importance Sampling Policy Optimization
von: Lu, Chengxuan, et al.
Veröffentlicht: (2026)
von: Lu, Chengxuan, et al.
Veröffentlicht: (2026)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
von: Li, Xuan, et al.
Veröffentlicht: (2026)
von: Li, Xuan, et al.
Veröffentlicht: (2026)
Rethinking Light Decoder-based Solvers for Vehicle Routing Problems
von: Huang, Ziwei, et al.
Veröffentlicht: (2025)
von: Huang, Ziwei, et al.
Veröffentlicht: (2025)
Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
Provable and Practical In-Context Policy Optimization for Self-Improvement
von: Yu, Tianrun, et al.
Veröffentlicht: (2026)
von: Yu, Tianrun, et al.
Veröffentlicht: (2026)
Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
von: Shen, Yixin
Veröffentlicht: (2025)
von: Shen, Yixin
Veröffentlicht: (2025)
Beyond Importance Sampling: Rejection-Gated Policy Optimization
von: Sun, Ziwu, et al.
Veröffentlicht: (2026)
von: Sun, Ziwu, et al.
Veröffentlicht: (2026)
Riemannian Batch Normalization: A Gyro Approach
von: Chen, Ziheng, et al.
Veröffentlicht: (2025)
von: Chen, Ziheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Knowledge Graph Embedding by Normalizing Flows
von: Xiao, Changyi, et al.
Veröffentlicht: (2024) -
Knowledge Graph Completion by Intermediate Variables Regularization
von: Xiao, Changyi, et al.
Veröffentlicht: (2025) -
Complex Logical Query Answering by Calibrating Knowledge Graph Completion Models
von: Xiao, Changyi, et al.
Veröffentlicht: (2024) -
Reinforcement Learning with Conditional Expectation Reward
von: Xiao, Changyi, et al.
Veröffentlicht: (2026) -
FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect Generalization
von: Hong, Shibo, et al.
Veröffentlicht: (2025)