Enhancing PPO with Trajectory-Aware Hybrid Policies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Qisai, Jiang, Zhanhong, Yang, Hsin-Jung, Khosravi, Mahsa, Waite, Joshua R., Sarkar, Soumik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
COOPO: Cyclic Offline-Online Policy Optimization Algorithm
von: Liu, Qisai, et al.
Veröffentlicht: (2026)
von: Liu, Qisai, et al.
Veröffentlicht: (2026)
Bidirectional Linear Recurrent Models for Sequence-Level Multisource Fusion
von: Liu, Qisai, et al.
Veröffentlicht: (2025)
von: Liu, Qisai, et al.
Veröffentlicht: (2025)
RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception
von: Waite, Joshua R., et al.
Veröffentlicht: (2025)
von: Waite, Joshua R., et al.
Veröffentlicht: (2025)
LexiSafe: Offline Safe Reinforcement Learning with Lexicographic Safety-Reward Hierarchy
von: Yang, Hsin-Jung, et al.
Veröffentlicht: (2026)
von: Yang, Hsin-Jung, et al.
Veröffentlicht: (2026)
Data-driven Kinematic Modeling in Soft Robots: System Identification and Uncertainty Quantification
von: Jiang, Zhanhong, et al.
Veröffentlicht: (2025)
von: Jiang, Zhanhong, et al.
Veröffentlicht: (2025)
TabQL: In-Context Q-Learning with Tabular Foundation Models
von: Liu, Qisai, et al.
Veröffentlicht: (2026)
von: Liu, Qisai, et al.
Veröffentlicht: (2026)
Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement Learning
von: Koirala, Prajwal, et al.
Veröffentlicht: (2024)
von: Koirala, Prajwal, et al.
Veröffentlicht: (2024)
DeCAF: Decentralized Consensus-And-Factorization for Low-Rank Adaptation of Foundation Models
von: Saadati, Nastaran, et al.
Veröffentlicht: (2025)
von: Saadati, Nastaran, et al.
Veröffentlicht: (2025)
FUSE: First-Order and Second-Order Unified SynthEsis in Stochastic Optimization
von: Jiang, Zhanhong, et al.
Veröffentlicht: (2025)
von: Jiang, Zhanhong, et al.
Veröffentlicht: (2025)
Incorporating System-level Safety Requirements in Perception Models via Reinforcement Learning
von: Fan, Weisi, et al.
Veröffentlicht: (2024)
von: Fan, Weisi, et al.
Veröffentlicht: (2024)
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning
von: Koirala, Prajwal, et al.
Veröffentlicht: (2024)
von: Koirala, Prajwal, et al.
Veröffentlicht: (2024)
Lighting-aware Unified Model for Instance Segmentation
von: Liu, Qisai, et al.
Veröffentlicht: (2026)
von: Liu, Qisai, et al.
Veröffentlicht: (2026)
DIMAT: Decentralized Iterative Merging-And-Training for Deep Learning Models
von: Saadati, Nastaran, et al.
Veröffentlicht: (2024)
von: Saadati, Nastaran, et al.
Veröffentlicht: (2024)
Neural CDEs as Correctors for Learned Time Series Models
von: Shahid, Muhammad Bilal, et al.
Veröffentlicht: (2025)
von: Shahid, Muhammad Bilal, et al.
Veröffentlicht: (2025)
Asynchronous Training Schemes in Distributed Learning with Time Delay
von: Wang, Haoxiang, et al.
Veröffentlicht: (2022)
von: Wang, Haoxiang, et al.
Veröffentlicht: (2022)
Balancing Utility and Privacy: Dynamically Private SGD with Random Projection
von: Jiang, Zhanhong, et al.
Veröffentlicht: (2025)
von: Jiang, Zhanhong, et al.
Veröffentlicht: (2025)
Zero-shot Sim-to-Real Transfer for Reinforcement Learning-based Visual Servoing of Soft Continuum Arms
von: Yang, Hsin-Jung, et al.
Veröffentlicht: (2025)
von: Yang, Hsin-Jung, et al.
Veröffentlicht: (2025)
ADKO: Agentic Decentralized Knowledge Optimization
von: Rillo, Lucas Nerone, et al.
Veröffentlicht: (2026)
von: Rillo, Lucas Nerone, et al.
Veröffentlicht: (2026)
Distributed Direct Preference Optimization
von: Jiang, Zhanhong
Veröffentlicht: (2026)
von: Jiang, Zhanhong
Veröffentlicht: (2026)
Optimizing Navigation And Chemical Application in Precision Agriculture With Deep Reinforcement Learning And Conditional Action Tree
von: Khosravi, Mahsa, et al.
Veröffentlicht: (2025)
von: Khosravi, Mahsa, et al.
Veröffentlicht: (2025)
STITCH: Surface reconstrucTion using Implicit neural representations with Topology Constraints and persistent Homology
von: Jignasu, Anushrut, et al.
Veröffentlicht: (2024)
von: Jignasu, Anushrut, et al.
Veröffentlicht: (2024)
ExO-PPO: an Extended Off-policy Proximal Policy Optimization Algorithm
von: Wang, Hanyong, et al.
Veröffentlicht: (2026)
von: Wang, Hanyong, et al.
Veröffentlicht: (2026)
Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative PPO
von: Jiang, Daniel R., et al.
Veröffentlicht: (2025)
von: Jiang, Daniel R., et al.
Veröffentlicht: (2025)
Decentralized Relaxed Smooth Optimization with Gradient Descent Methods
von: Jiang, Zhanhong, et al.
Veröffentlicht: (2025)
von: Jiang, Zhanhong, et al.
Veröffentlicht: (2025)
Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
von: Li, Junbo, et al.
Veröffentlicht: (2025)
von: Li, Junbo, et al.
Veröffentlicht: (2025)
Eval-PPO: Building an Efficient Threat Evaluator Using Proximal Policy Optimization
von: Sun, Wuzhou, et al.
Veröffentlicht: (2025)
von: Sun, Wuzhou, et al.
Veröffentlicht: (2025)
Directional-Clamp PPO
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
BinaryPPO: Efficient Policy Optimization for Binary Classification
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2026)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2026)
Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training
von: Gong, Xue, et al.
Veröffentlicht: (2026)
von: Gong, Xue, et al.
Veröffentlicht: (2026)
AgGym: An agricultural biotic stress simulation environment for ultra-precision management planning
von: Khosravi, Mahsa, et al.
Veröffentlicht: (2024)
von: Khosravi, Mahsa, et al.
Veröffentlicht: (2024)
CIM-PPO:Proximal Policy Optimization with Liu-Correntropy Induced Metric
von: Guo, Yunxiao, et al.
Veröffentlicht: (2021)
von: Guo, Yunxiao, et al.
Veröffentlicht: (2021)
Traffic State Estimation from Vehicle Trajectories with Anisotropic Gaussian Processes
von: Wu, Fan, et al.
Veröffentlicht: (2023)
von: Wu, Fan, et al.
Veröffentlicht: (2023)
Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement Learning
von: Daley, Brett, et al.
Veröffentlicht: (2023)
von: Daley, Brett, et al.
Veröffentlicht: (2023)
CATP: Context-Aware Trajectory Prediction with Competition Symbiosis
von: Wu, Jiang, et al.
Veröffentlicht: (2024)
von: Wu, Jiang, et al.
Veröffentlicht: (2024)
STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
von: Chen, Yuhan, et al.
Veröffentlicht: (2025)
von: Chen, Yuhan, et al.
Veröffentlicht: (2025)
AM-PPO: (Advantage) Alpha-Modulation with Proximal Policy Optimization
von: Sane, Soham
Veröffentlicht: (2025)
von: Sane, Soham
Veröffentlicht: (2025)
Sampling Complexity of TD and PPO in RKHS
von: Zou, Lu, et al.
Veröffentlicht: (2025)
von: Zou, Lu, et al.
Veröffentlicht: (2025)
Enhancing Generalization via Sharpness-Aware Trajectory Matching for Dataset Condensation
von: Gao, Boyan, et al.
Veröffentlicht: (2025)
von: Gao, Boyan, et al.
Veröffentlicht: (2025)
PPO-Based Hybrid Optimization for RIS-Assisted Semantic Vehicular Edge Computing
von: Feng, Wei, et al.
Veröffentlicht: (2026)
von: Feng, Wei, et al.
Veröffentlicht: (2026)
Non-Asymptotic Global Convergence of PPO-Clip
von: Liu, Yin, et al.
Veröffentlicht: (2025)
von: Liu, Yin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
COOPO: Cyclic Offline-Online Policy Optimization Algorithm
von: Liu, Qisai, et al.
Veröffentlicht: (2026) -
Bidirectional Linear Recurrent Models for Sequence-Level Multisource Fusion
von: Liu, Qisai, et al.
Veröffentlicht: (2025) -
RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception
von: Waite, Joshua R., et al.
Veröffentlicht: (2025) -
LexiSafe: Offline Safe Reinforcement Learning with Lexicographic Safety-Reward Hierarchy
von: Yang, Hsin-Jung, et al.
Veröffentlicht: (2026) -
Data-driven Kinematic Modeling in Soft Robots: System Identification and Uncertainty Quantification
von: Jiang, Zhanhong, et al.
Veröffentlicht: (2025)