Revisiting Discrete Soft Actor-Critic
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Haibin, Wei, Tong, Lin, Zichuan, li, junyou, Xing, Junliang, Shi, Yuanchun, Shen, Li, Yu, Chao, Ye, Deheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training
von: Wei, Tong, et al.
Veröffentlicht: (2025)
von: Wei, Tong, et al.
Veröffentlicht: (2025)
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
von: Wei, Tong, et al.
Veröffentlicht: (2025)
von: Wei, Tong, et al.
Veröffentlicht: (2025)
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
von: Asad, Reza, et al.
Veröffentlicht: (2025)
von: Asad, Reza, et al.
Veröffentlicht: (2025)
HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation
von: Zhao, He, et al.
Veröffentlicht: (2026)
von: Zhao, He, et al.
Veröffentlicht: (2026)
Learning Versatile Skills with Curriculum Masking
von: Tang, Yao, et al.
Veröffentlicht: (2024)
von: Tang, Yao, et al.
Veröffentlicht: (2024)
Distributional Soft Actor-Critic with Diffusion Policy
von: Liu, Tong, et al.
Veröffentlicht: (2025)
von: Liu, Tong, et al.
Veröffentlicht: (2025)
CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasks
von: Chai, Qi, et al.
Veröffentlicht: (2025)
von: Chai, Qi, et al.
Veröffentlicht: (2025)
Average-Reward Soft Actor-Critic
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
von: Lin, Zichuan, et al.
Veröffentlicht: (2025)
von: Lin, Zichuan, et al.
Veröffentlicht: (2025)
Rethinking Soft Actor-Critic in High-Dimensional Action Spaces: The Cost of Ignoring Distribution Shift
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
von: Ma, Xiaoteng, et al.
Veröffentlicht: (2020)
von: Ma, Xiaoteng, et al.
Veröffentlicht: (2020)
SACn: Soft Actor-Critic with n-step Returns
von: Łyskawa, Jakub, et al.
Veröffentlicht: (2025)
von: Łyskawa, Jakub, et al.
Veröffentlicht: (2025)
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
von: He, Jiamin, et al.
Veröffentlicht: (2026)
von: He, Jiamin, et al.
Veröffentlicht: (2026)
PIPCFR: Pseudo-outcome Imputation with Post-treatment Variables for Individual Treatment Effect Estimation
von: Lin, Zichuan, et al.
Veröffentlicht: (2025)
von: Lin, Zichuan, et al.
Veröffentlicht: (2025)
DeCoDe: Defer-and-Complement Decision-Making via Decoupled Concept Bottleneck Models
von: He, Chengbo, et al.
Veröffentlicht: (2025)
von: He, Chengbo, et al.
Veröffentlicht: (2025)
Decorrelated Soft Actor-Critic for Efficient Deep Reinforcement Learning
von: Küçükoğlu, Burcu, et al.
Veröffentlicht: (2025)
von: Küçükoğlu, Burcu, et al.
Veröffentlicht: (2025)
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning
von: Zhang, Yixian, et al.
Veröffentlicht: (2025)
von: Zhang, Yixian, et al.
Veröffentlicht: (2025)
Debiased Model-based Representations for Sample-efficient Continuous Control
von: Lyu, Jiafei, et al.
Veröffentlicht: (2026)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2026)
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
von: Yang, Kai, et al.
Veröffentlicht: (2025)
von: Yang, Kai, et al.
Veröffentlicht: (2025)
DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty
von: Cui, Mingxuan, et al.
Veröffentlicht: (2025)
von: Cui, Mingxuan, et al.
Veröffentlicht: (2025)
AI Olympics challenge with Evolutionary Soft Actor Critic
von: Calì, Marco, et al.
Veröffentlicht: (2024)
von: Calì, Marco, et al.
Veröffentlicht: (2024)
Causal Policy Learning in Reinforcement Learning: Backdoor-Adjusted Soft Actor-Critic
von: Vo, Thanh Vinh, et al.
Veröffentlicht: (2025)
von: Vo, Thanh Vinh, et al.
Veröffentlicht: (2025)
CTSAC: Curriculum-Based Transformer Soft Actor-Critic for Goal-Oriented Robot Exploration
von: Yang, Chunyu, et al.
Veröffentlicht: (2025)
von: Yang, Chunyu, et al.
Veröffentlicht: (2025)
Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration
von: Bai, Qinxun, et al.
Veröffentlicht: (2025)
von: Bai, Qinxun, et al.
Veröffentlicht: (2025)
Diffusion Actor-Critic with Entropy Regulator
von: Wang, Yinuo, et al.
Veröffentlicht: (2024)
von: Wang, Yinuo, et al.
Veröffentlicht: (2024)
Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion
von: Sabatini, Gianluca, et al.
Veröffentlicht: (2026)
von: Sabatini, Gianluca, et al.
Veröffentlicht: (2026)
Towards Safe Maneuvering of Double-Ackermann-Steering Robots with a Soft Actor-Critic Framework
von: Deflesselle, Kohio, et al.
Veröffentlicht: (2025)
von: Deflesselle, Kohio, et al.
Veröffentlicht: (2025)
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
von: Yang, Tong, et al.
Veröffentlicht: (2023)
von: Yang, Tong, et al.
Veröffentlicht: (2023)
Actor-Critics Can Achieve Optimal Sample Efficiency
von: Tan, Kevin, et al.
Veröffentlicht: (2025)
von: Tan, Kevin, et al.
Veröffentlicht: (2025)
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
von: Lu, Zhicong, et al.
Veröffentlicht: (2026)
von: Lu, Zhicong, et al.
Veröffentlicht: (2026)
Reinforcement Learning Position Control of a Quadrotor Using Soft Actor-Critic (SAC)
von: Mahran, Youssef, et al.
Veröffentlicht: (2025)
von: Mahran, Youssef, et al.
Veröffentlicht: (2025)
SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling
von: Gaven, Loris, et al.
Veröffentlicht: (2024)
von: Gaven, Loris, et al.
Veröffentlicht: (2024)
ProAct: Agentic Lookahead in Interactive Environments
von: Yu, Yangbin, et al.
Veröffentlicht: (2026)
von: Yu, Yangbin, et al.
Veröffentlicht: (2026)
Application of Soft Actor-Critic Algorithms in Optimizing Wastewater Treatment with Time Delays Integration
von: Mohammadi, Esmaeel, et al.
Veröffentlicht: (2024)
von: Mohammadi, Esmaeel, et al.
Veröffentlicht: (2024)
Pareto Actor-Critic for Communication and Computation Co-Optimization in Non-Cooperative Federated Learning Services
von: Tan, Renxuan, et al.
Veröffentlicht: (2025)
von: Tan, Renxuan, et al.
Veröffentlicht: (2025)
Proximal Action Replacement for Behavior Cloning Actor-Critic in Offline Reinforcement Learning
von: Dong, Jinzong, et al.
Veröffentlicht: (2026)
von: Dong, Jinzong, et al.
Veröffentlicht: (2026)
K^2-Agent: Co-Evolving Know-What and Know-How for Hierarchical Mobile Device Control
von: Wu, Zhe, et al.
Veröffentlicht: (2026)
von: Wu, Zhe, et al.
Veröffentlicht: (2026)
Asymmetric Actor-Critic for Multi-turn LLM Agents
von: Jiang, Shuli, et al.
Veröffentlicht: (2026)
von: Jiang, Shuli, et al.
Veröffentlicht: (2026)
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
von: Zhang, Feihong, et al.
Veröffentlicht: (2025)
von: Zhang, Feihong, et al.
Veröffentlicht: (2025)
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training
von: Wei, Tong, et al.
Veröffentlicht: (2025) -
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
von: Wei, Tong, et al.
Veröffentlicht: (2025) -
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
von: Asad, Reza, et al.
Veröffentlicht: (2025) -
HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation
von: Zhao, He, et al.
Veröffentlicht: (2026) -
Learning Versatile Skills with Curriculum Masking
von: Tang, Yao, et al.
Veröffentlicht: (2024)