Gespeichert in:
| 1. Verfasser: | Spigler, Giacomo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2407.15134 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning
von: Spigler, Giacomo
Veröffentlicht: (2026)
von: Spigler, Giacomo
Veröffentlicht: (2026)
Predicting Depression and Anxiety Risk in Dutch Neighborhoods from Street-View Images
von: Khodorivsko, Nin, et al.
Veröffentlicht: (2024)
von: Khodorivsko, Nin, et al.
Veröffentlicht: (2024)
Imitation of human motion achieves natural head movements for humanoid robots in an active-speaker detection task
von: Ding, Bosong, et al.
Veröffentlicht: (2024)
von: Ding, Bosong, et al.
Veröffentlicht: (2024)
Reparameterization Proximal Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)
Beyond the Boundaries of Proximal Policy Optimization
von: Tan, Charlie B., et al.
Veröffentlicht: (2024)
von: Tan, Charlie B., et al.
Veröffentlicht: (2024)
Proximal Policy Optimization with Adaptive Exploration
von: Lixandru, Andrei
Veröffentlicht: (2024)
von: Lixandru, Andrei
Veröffentlicht: (2024)
Complexity-Regularized Proximal Policy Optimization
von: Serfilippi, Luca, et al.
Veröffentlicht: (2025)
von: Serfilippi, Luca, et al.
Veröffentlicht: (2025)
KIPPO: Koopman-Inspired Proximal Policy Optimization
von: Cozma, Andrei, et al.
Veröffentlicht: (2025)
von: Cozma, Andrei, et al.
Veröffentlicht: (2025)
ESPO: Early-Stopping Proximal Policy Optimization
von: Li, Zihang, et al.
Veröffentlicht: (2026)
von: Li, Zihang, et al.
Veröffentlicht: (2026)
Learning Branching Policies for MILPs with Proximal Policy Optimization
von: Mhamed, Abdelouahed Ben, et al.
Veröffentlicht: (2025)
von: Mhamed, Abdelouahed Ben, et al.
Veröffentlicht: (2025)
Extreme Region Policy Distillation
von: Chen, Changyu, et al.
Veröffentlicht: (2026)
von: Chen, Changyu, et al.
Veröffentlicht: (2026)
Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement Learning
von: Batra, Sumeet, et al.
Veröffentlicht: (2023)
von: Batra, Sumeet, et al.
Veröffentlicht: (2023)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
von: Liu, Jiashun, et al.
Veröffentlicht: (2025)
von: Liu, Jiashun, et al.
Veröffentlicht: (2025)
CIM-PPO:Proximal Policy Optimization with Liu-Correntropy Induced Metric
von: Guo, Yunxiao, et al.
Veröffentlicht: (2021)
von: Guo, Yunxiao, et al.
Veröffentlicht: (2021)
A dynamical clipping approach with task feedback for Proximal Policy Optimization
von: Zhang, Ziqi, et al.
Veröffentlicht: (2023)
von: Zhang, Ziqi, et al.
Veröffentlicht: (2023)
PROMA: Projected Microbatch Accumulation for Reference-Free Proximal Policy Updates
von: Abrahamsen, Nilin
Veröffentlicht: (2026)
von: Abrahamsen, Nilin
Veröffentlicht: (2026)
Efficient Deep Reinforcement Learning with Predictive Processing Proximal Policy Optimization
von: Küçükoğlu, Burcu, et al.
Veröffentlicht: (2022)
von: Küçükoğlu, Burcu, et al.
Veröffentlicht: (2022)
Online Policy Distillation with Decision-Attention
von: Yu, Xinqiang, et al.
Veröffentlicht: (2024)
von: Yu, Xinqiang, et al.
Veröffentlicht: (2024)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
von: Ding, Ken
Veröffentlicht: (2026)
von: Ding, Ken
Veröffentlicht: (2026)
TIP: Token Importance in On-Policy Distillation
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
ExO-PPO: an Extended Off-policy Proximal Policy Optimization Algorithm
von: Wang, Hanyong, et al.
Veröffentlicht: (2026)
von: Wang, Hanyong, et al.
Veröffentlicht: (2026)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
Trust-Region Behavior Blending for On-Policy Distillation
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs
von: Kohler, Hector, et al.
Veröffentlicht: (2025)
von: Kohler, Hector, et al.
Veröffentlicht: (2025)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
von: Liang, Kun, et al.
Veröffentlicht: (2026)
von: Liang, Kun, et al.
Veröffentlicht: (2026)
Stable On-Policy Distillation through Adaptive Target Reformulation
von: Jang, Ijun, et al.
Veröffentlicht: (2026)
von: Jang, Ijun, et al.
Veröffentlicht: (2026)
Interpretable Policy Distillation for Power Grid Topology Control
von: Dmitruka, Aleksandra, et al.
Veröffentlicht: (2026)
von: Dmitruka, Aleksandra, et al.
Veröffentlicht: (2026)
Explainable RL Policies by Distilling to Locally-Specialized Linear Policies with Voronoi State Partitioning
von: Deproost, Senne, et al.
Veröffentlicht: (2025)
von: Deproost, Senne, et al.
Veröffentlicht: (2025)
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards
von: Ahmad, Ahmad, et al.
Veröffentlicht: (2024)
von: Ahmad, Ahmad, et al.
Veröffentlicht: (2024)
Variational Distillation of Diffusion Policies into Mixture of Experts
von: Zhou, Hongyi, et al.
Veröffentlicht: (2024)
von: Zhou, Hongyi, et al.
Veröffentlicht: (2024)
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
von: Jia, Nan, et al.
Veröffentlicht: (2026)
von: Jia, Nan, et al.
Veröffentlicht: (2026)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
von: Armandpour, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Armandpour, Mohammadreza, et al.
Veröffentlicht: (2026)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
von: Weltevrede, Max, et al.
Veröffentlicht: (2025)
von: Weltevrede, Max, et al.
Veröffentlicht: (2025)
A Deep Reinforcement Learning Approach to Battery Management in Dairy Farming via Proximal Policy Optimization
von: Ali, Nawazish, et al.
Veröffentlicht: (2024)
von: Ali, Nawazish, et al.
Veröffentlicht: (2024)
Proximity Matters: Local Proximity Enhanced Balancing for Treatment Effect Estimation
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning
von: Spigler, Giacomo
Veröffentlicht: (2026) -
Predicting Depression and Anxiety Risk in Dutch Neighborhoods from Street-View Images
von: Khodorivsko, Nin, et al.
Veröffentlicht: (2024) -
Imitation of human motion achieves natural head movements for humanoid robots in an active-speaker detection task
von: Ding, Bosong, et al.
Veröffentlicht: (2024) -
Reparameterization Proximal Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2025) -
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)