Average-Reward Soft Actor-Critic
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Adamczyk, Jacob, Makarenko, Volodymyr, Tiomkin, Stas, Kulkarni, Rahul V. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bootstrapped Reward Shaping
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
EVAL: EigenVector-based Average-reward Learning
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
Boosting Soft Q-Learning by Bounding
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2024)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2024)
Maximum Entropy Exploration Without the Rollouts
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026)
Thermodynamics of Reinforcement Learning Curricula
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Efficient $Q$-Learning and Actor-Critic Methods for Robust Average Reward Reinforcement Learning
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Multi-Resolution Diffusion for Privacy-Sensitive Recommender Systems
von: Lilienthal, Derek, et al.
Veröffentlicht: (2023)
von: Lilienthal, Derek, et al.
Veröffentlicht: (2023)
Revisiting Discrete Soft Actor-Critic
von: Zhou, Haibin, et al.
Veröffentlicht: (2022)
von: Zhou, Haibin, et al.
Veröffentlicht: (2022)
SuPLE: Robot Learning with Lyapunov Rewards
von: Nguyen, Phu, et al.
Veröffentlicht: (2024)
von: Nguyen, Phu, et al.
Veröffentlicht: (2024)
Distributional Soft Actor-Critic with Diffusion Policy
von: Liu, Tong, et al.
Veröffentlicht: (2025)
von: Liu, Tong, et al.
Veröffentlicht: (2025)
Exploration Behavior of Untrained Policies
von: Adamczyk, Jacob
Veröffentlicht: (2025)
von: Adamczyk, Jacob
Veröffentlicht: (2025)
Inferring Transition Dynamics from Value Functions
von: Adamczyk, Jacob
Veröffentlicht: (2025)
von: Adamczyk, Jacob
Veröffentlicht: (2025)
SACn: Soft Actor-Critic with n-step Returns
von: Łyskawa, Jakub, et al.
Veröffentlicht: (2025)
von: Łyskawa, Jakub, et al.
Veröffentlicht: (2025)
Emergence of Physical Intelligence via Controllable Information Production
von: Shah, Tristan, et al.
Veröffentlicht: (2026)
von: Shah, Tristan, et al.
Veröffentlicht: (2026)
Learning telic-controllable state representations
von: Amir, Nadav, et al.
Veröffentlicht: (2024)
von: Amir, Nadav, et al.
Veröffentlicht: (2024)
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
von: Kar, Avik, et al.
Veröffentlicht: (2024)
von: Kar, Avik, et al.
Veröffentlicht: (2024)
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
von: Asad, Reza, et al.
Veröffentlicht: (2025)
von: Asad, Reza, et al.
Veröffentlicht: (2025)
Decorrelated Soft Actor-Critic for Efficient Deep Reinforcement Learning
von: Küçükoğlu, Burcu, et al.
Veröffentlicht: (2025)
von: Küçükoğlu, Burcu, et al.
Veröffentlicht: (2025)
DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
von: Ma, Xiaoteng, et al.
Veröffentlicht: (2020)
von: Ma, Xiaoteng, et al.
Veröffentlicht: (2020)
Causal Policy Learning in Reinforcement Learning: Backdoor-Adjusted Soft Actor-Critic
von: Vo, Thanh Vinh, et al.
Veröffentlicht: (2025)
von: Vo, Thanh Vinh, et al.
Veröffentlicht: (2025)
Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion
von: Sabatini, Gianluca, et al.
Veröffentlicht: (2026)
von: Sabatini, Gianluca, et al.
Veröffentlicht: (2026)
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning
von: Zhang, Yixian, et al.
Veröffentlicht: (2025)
von: Zhang, Yixian, et al.
Veröffentlicht: (2025)
SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling
von: Gaven, Loris, et al.
Veröffentlicht: (2024)
von: Gaven, Loris, et al.
Veröffentlicht: (2024)
Rethinking Soft Actor-Critic in High-Dimensional Action Spaces: The Cost of Ignoring Distribution Shift
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
AI Olympics challenge with Evolutionary Soft Actor Critic
von: Calì, Marco, et al.
Veröffentlicht: (2024)
von: Calì, Marco, et al.
Veröffentlicht: (2024)
Reinforcement Learning Position Control of a Quadrotor Using Soft Actor-Critic (SAC)
von: Mahran, Youssef, et al.
Veröffentlicht: (2025)
von: Mahran, Youssef, et al.
Veröffentlicht: (2025)
Value Improved Actor Critic Algorithms
von: Oren, Yaniv, et al.
Veröffentlicht: (2024)
von: Oren, Yaniv, et al.
Veröffentlicht: (2024)
Diffusion Actor-Critic with Entropy Regulator
von: Wang, Yinuo, et al.
Veröffentlicht: (2024)
von: Wang, Yinuo, et al.
Veröffentlicht: (2024)
DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty
von: Cui, Mingxuan, et al.
Veröffentlicht: (2025)
von: Cui, Mingxuan, et al.
Veröffentlicht: (2025)
Application of Soft Actor-Critic Algorithms in Optimizing Wastewater Treatment with Time Delays Integration
von: Mohammadi, Esmaeel, et al.
Veröffentlicht: (2024)
von: Mohammadi, Esmaeel, et al.
Veröffentlicht: (2024)
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
von: Zhang, Feihong, et al.
Veröffentlicht: (2025)
von: Zhang, Feihong, et al.
Veröffentlicht: (2025)
Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
von: Garcin, Samuel, et al.
Veröffentlicht: (2025)
von: Garcin, Samuel, et al.
Veröffentlicht: (2025)
Flow Actor-Critic for Offline Reinforcement Learning
von: Chae, Jongseong, et al.
Veröffentlicht: (2026)
von: Chae, Jongseong, et al.
Veröffentlicht: (2026)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
von: Manivannan, Sanjeev, et al.
Veröffentlicht: (2026)
von: Manivannan, Sanjeev, et al.
Veröffentlicht: (2026)
WARP: On the Benefits of Weight Averaged Rewarded Policies
von: Ramé, Alexandre, et al.
Veröffentlicht: (2024)
von: Ramé, Alexandre, et al.
Veröffentlicht: (2024)
Goals and the Structure of Experience
von: Amir, Nadav, et al.
Veröffentlicht: (2025)
von: Amir, Nadav, et al.
Veröffentlicht: (2025)
Decentralized Traffic Flow Optimization Through Intrinsic Motivation
von: Papala, Himaja, et al.
Veröffentlicht: (2025)
von: Papala, Himaja, et al.
Veröffentlicht: (2025)
Relational Object-Centric Actor-Critic
von: Ugadiarov, Leonid, et al.
Veröffentlicht: (2023)
von: Ugadiarov, Leonid, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Bootstrapped Reward Shaping
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025) -
EVAL: EigenVector-based Average-reward Learning
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025) -
Boosting Soft Q-Learning by Bounding
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2024) -
Maximum Entropy Exploration Without the Rollouts
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026) -
Thermodynamics of Reinforcement Learning Curricula
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026)