Directional-Clamp PPO
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Karpel, Gilad, Zhou, Ruida, Sabach, Shoham, Ghavamzadeh, Mohammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
von: Viano, Luca, et al.
Veröffentlicht: (2026)
von: Viano, Luca, et al.
Veröffentlicht: (2026)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
von: Pipano, Idan, et al.
Veröffentlicht: (2026)
von: Pipano, Idan, et al.
Veröffentlicht: (2026)
C2-DPO: Constrained Controlled Direct Preference Optimization
von: Asadi, Kavosh, et al.
Veröffentlicht: (2025)
von: Asadi, Kavosh, et al.
Veröffentlicht: (2025)
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026)
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026)
Learning the Target Network in Function Space
von: Asadi, Kavosh, et al.
Veröffentlicht: (2024)
von: Asadi, Kavosh, et al.
Veröffentlicht: (2024)
FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain
von: Deb, Rohan, et al.
Veröffentlicht: (2025)
von: Deb, Rohan, et al.
Veröffentlicht: (2025)
Krylov Cubic Regularized Newton: A Subspace Second-Order Method with Dimension-Free Convergence Rate
von: Jiang, Ruichen, et al.
Veröffentlicht: (2024)
von: Jiang, Ruichen, et al.
Veröffentlicht: (2024)
DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2026)
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2026)
Optimal L2 Regularization in High-dimensional Continual Linear Regression
von: Karpel, Gilad, et al.
Veröffentlicht: (2026)
von: Karpel, Gilad, et al.
Veröffentlicht: (2026)
Bayesian Regret Minimization in Offline Bandits
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
Bayesian policy gradient and actor-critic algorithms
von: Ghavamzadeh, Mohammad, et al.
Veröffentlicht: (2026)
von: Ghavamzadeh, Mohammad, et al.
Veröffentlicht: (2026)
Contextual Bandits with Stage-wise Constraints
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
von: Thekumparampil, Kiran Koshy, et al.
Veröffentlicht: (2024)
von: Thekumparampil, Kiran Koshy, et al.
Veröffentlicht: (2024)
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
von: Ozkara, Kaan, et al.
Veröffentlicht: (2024)
von: Ozkara, Kaan, et al.
Veröffentlicht: (2024)
Conservative Contextual Bandits: Beyond Linear Representations
von: Deb, Rohan, et al.
Veröffentlicht: (2024)
von: Deb, Rohan, et al.
Veröffentlicht: (2024)
TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained Models
von: Liu, Zuxin, et al.
Veröffentlicht: (2023)
von: Liu, Zuxin, et al.
Veröffentlicht: (2023)
Maximum Entropy Semi-Supervised Inverse Reinforcement Learning
von: Audiffren, Julien, et al.
Veröffentlicht: (2026)
von: Audiffren, Julien, et al.
Veröffentlicht: (2026)
A Proximal Operator for Inducing 2:4-Sparsity
von: Kübler, Jonas M, et al.
Veröffentlicht: (2025)
von: Kübler, Jonas M, et al.
Veröffentlicht: (2025)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
von: Liu, Hongyi, et al.
Veröffentlicht: (2025)
von: Liu, Hongyi, et al.
Veröffentlicht: (2025)
Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
von: Li, Junbo, et al.
Veröffentlicht: (2025)
von: Li, Junbo, et al.
Veröffentlicht: (2025)
RASR: Risk-Averse Soft-Robust MDPs with EVaR and Entropic Risk
von: Hau, Jia Lin, et al.
Veröffentlicht: (2022)
von: Hau, Jia Lin, et al.
Veröffentlicht: (2022)
Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis
von: Hau, Jia Lin, et al.
Veröffentlicht: (2024)
von: Hau, Jia Lin, et al.
Veröffentlicht: (2024)
SPIRE: Conditional Personalization for Federated Diffusion Generative Models
von: Ozkara, Kaan, et al.
Veröffentlicht: (2025)
von: Ozkara, Kaan, et al.
Veröffentlicht: (2025)
Correlational Lagrangian Schrödinger Bridge: Learning Dynamics with Population-Level Regularization
von: You, Yuning, et al.
Veröffentlicht: (2024)
von: You, Yuning, et al.
Veröffentlicht: (2024)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
Sampling Complexity of TD and PPO in RKHS
von: Zou, Lu, et al.
Veröffentlicht: (2025)
von: Zou, Lu, et al.
Veröffentlicht: (2025)
Differentially private ratio statistics
von: Shoham, Tomer, et al.
Veröffentlicht: (2025)
von: Shoham, Tomer, et al.
Veröffentlicht: (2025)
Unbiased Stochastic Optimization for Gaussian Processes on Finite Dimensional RKHS
von: Shoham, Neta, et al.
Veröffentlicht: (2025)
von: Shoham, Neta, et al.
Veröffentlicht: (2025)
An Information-Theoretic Approach to Understanding Transformers' In-Context Learning of Variable-Order Markov Chains
von: Zhou, Ruida, et al.
Veröffentlicht: (2024)
von: Zhou, Ruida, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning with Enhanced PPO for Safe Mobile Robot Navigation
von: Taheri, Hamid, et al.
Veröffentlicht: (2024)
von: Taheri, Hamid, et al.
Veröffentlicht: (2024)
PPO in the Fisher-Rao geometry
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2025)
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2025)
Enhancing PPO with Trajectory-Aware Hybrid Policies
von: Liu, Qisai, et al.
Veröffentlicht: (2025)
von: Liu, Qisai, et al.
Veröffentlicht: (2025)
From Function to Distribution Modeling: A PAC-Generative Approach to Offline Optimization
von: Zhang, Qiang, et al.
Veröffentlicht: (2024)
von: Zhang, Qiang, et al.
Veröffentlicht: (2024)
ADEPT: Hierarchical Bayes Approach to Personalized Federated Unsupervised Learning
von: Ozkara, Kaan, et al.
Veröffentlicht: (2024)
von: Ozkara, Kaan, et al.
Veröffentlicht: (2024)
A Unified Python Framework for Direct PPO-based Control of AHUs with Economizer Logic and CO2-Constrained Ventilation
von: Damavandi, Erfan Haghighat, et al.
Veröffentlicht: (2026)
von: Damavandi, Erfan Haghighat, et al.
Veröffentlicht: (2026)
On the optimal regret of collaborative personalized linear bandits
von: Huang, Bruce, et al.
Veröffentlicht: (2025)
von: Huang, Bruce, et al.
Veröffentlicht: (2025)
Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models
von: Kim, Kyuyoung, et al.
Veröffentlicht: (2024)
von: Kim, Kyuyoung, et al.
Veröffentlicht: (2024)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
Dynamic FISTA for Convex Composite Bi-Level Optimization
von: Merchav, Roey, et al.
Veröffentlicht: (2024)
von: Merchav, Roey, et al.
Veröffentlicht: (2024)
Neural Clamping: Joint Input Perturbation and Temperature Scaling for Neural Network Calibration
von: Tang, Yung-Chen, et al.
Veröffentlicht: (2022)
von: Tang, Yung-Chen, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
von: Viano, Luca, et al.
Veröffentlicht: (2026) -
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
von: Pipano, Idan, et al.
Veröffentlicht: (2026) -
C2-DPO: Constrained Controlled Direct Preference Optimization
von: Asadi, Kavosh, et al.
Veröffentlicht: (2025) -
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026) -
Learning the Target Network in Function Space
von: Asadi, Kavosh, et al.
Veröffentlicht: (2024)