ACE : Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
Fuente:
arXiv
Saved in:
| Main Authors: | Ji, Tianying, Liang, Yongyuan, Zeng, Yan, Luo, Yu, Xu, Guowei, Guo, Jiawei, Zheng, Ruijie, Huang, Furong, Sun, Fuchun, Xu, Huazhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-Critic
by: Ji, Tianying, et al.
Published: (2023)
by: Ji, Tianying, et al.
Published: (2023)
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted Diffusion
by: Liang, Yongyuan, et al.
Published: (2024)
by: Liang, Yongyuan, et al.
Published: (2024)
Soft Actor-Critic with Beta Policy via Implicit Reparameterization Gradients
by: Della Libera, Luca
Published: (2024)
by: Della Libera, Luca
Published: (2024)
Enhancing Multi-Agent Collaboration with Attention-Based Actor-Critic Policies
by: Belinchon, Hugo Garrido-Lestache, et al.
Published: (2025)
by: Belinchon, Hugo Garrido-Lestache, et al.
Published: (2025)
DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization
by: Xu, Guowei, et al.
Published: (2023)
by: Xu, Guowei, et al.
Published: (2023)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
by: Ged, François, et al.
Published: (2023)
by: Ged, François, et al.
Published: (2023)
Bidirectional-Reachable Hierarchical Reinforcement Learning with Mutually Responsive Policies
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
by: He, Jiamin, et al.
Published: (2026)
by: He, Jiamin, et al.
Published: (2026)
Premier-TACO is a Few-Shot Policy Learner: Pretraining Multitask Representation via Temporal Action-Driven Contrastive Loss
by: Zheng, Ruijie, et al.
Published: (2024)
by: Zheng, Ruijie, et al.
Published: (2024)
Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization
by: Cooper, Patrick, et al.
Published: (2026)
by: Cooper, Patrick, et al.
Published: (2026)
Actor-Critic Model Predictive Control: Differentiable Optimization meets Reinforcement Learning for Agile Flight
by: Romero, Angel, et al.
Published: (2023)
by: Romero, Angel, et al.
Published: (2023)
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs
by: Tang, Guowei
Published: (2026)
by: Tang, Guowei
Published: (2026)
Measuring Reasoning Utility in LLMs via Conditional Entropy Reduction
by: Guo, Xu
Published: (2025)
by: Guo, Xu
Published: (2025)
Privacy-Preserving Explainable AIoT Application via SHAP Entropy Regularization
by: Sharma, Dilli Prasad, et al.
Published: (2025)
by: Sharma, Dilli Prasad, et al.
Published: (2025)
TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
by: Zheng, Ruijie, et al.
Published: (2024)
by: Zheng, Ruijie, et al.
Published: (2024)
Detecting Subtle Differences between Human and Model Languages Using Spectrum of Relative Likelihood
by: Xu, Yang, et al.
Published: (2024)
by: Xu, Yang, et al.
Published: (2024)
Entropy Causal Graphs for Multivariate Time Series Anomaly Detection
by: Febrinanto, Falih Gozi, et al.
Published: (2023)
by: Febrinanto, Falih Gozi, et al.
Published: (2023)
Entropy Aware Message Passing in Graph Neural Networks
by: Nazari, Philipp, et al.
Published: (2024)
by: Nazari, Philipp, et al.
Published: (2024)
RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation
by: Chen, Junting, et al.
Published: (2024)
by: Chen, Junting, et al.
Published: (2024)
DemoSpeedup: Accelerating Visuomotor Policies via Entropy-Guided Demonstration Acceleration
by: Guo, Lingxiao, et al.
Published: (2025)
by: Guo, Lingxiao, et al.
Published: (2025)
Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search
by: Ma, Shubin, et al.
Published: (2025)
by: Ma, Shubin, et al.
Published: (2025)
ZAPS-DA: Zero-Phase Action Policy Smoothing with Decoupled Actor for Continuous Control in Reinforcement Learning
by: Shamass, Faiq
Published: (2026)
by: Shamass, Faiq
Published: (2026)
Physics-Informed Policy Optimization via Analytic Dynamics Regularization
by: Chandra, Namai, et al.
Published: (2026)
by: Chandra, Namai, et al.
Published: (2026)
VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
RoboGolf: Mastering Real-World Minigolf with a Reflective Multi-Modality Vision-Language Model
by: Zhou, Hantao, et al.
Published: (2024)
by: Zhou, Hantao, et al.
Published: (2024)
Evaluating and Learning Robust Bandit Policies Under Uncertain Causal Mechanisms
by: Avery, Katherine, et al.
Published: (2025)
by: Avery, Katherine, et al.
Published: (2025)
Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration
by: Bai, Qinxun, et al.
Published: (2025)
by: Bai, Qinxun, et al.
Published: (2025)
Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
by: Kang, Zilin, et al.
Published: (2025)
by: Kang, Zilin, et al.
Published: (2025)
Logging Policy Design for Off-Policy Evaluation
by: Douglas, Connor, et al.
Published: (2026)
by: Douglas, Connor, et al.
Published: (2026)
Dynamic Policy Induction for Adaptive Prompt Optimization: Bridging the Efficiency-Accuracy Gap via Lightweight Reinforcement Learning
by: Xu, Jiexi
Published: (2025)
by: Xu, Jiexi
Published: (2025)
Off-Policy Actor-Critic with Sigmoid-Bounded Entropy for Real-World Robot Learning
by: Wu, Xiefeng, et al.
Published: (2026)
by: Wu, Xiefeng, et al.
Published: (2026)
Refined Analysis of Entropy-Regularized Actor-Critic
by: Labbi, Safwan, et al.
Published: (2026)
by: Labbi, Safwan, et al.
Published: (2026)
RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
by: Ge, Junyao, et al.
Published: (2024)
by: Ge, Junyao, et al.
Published: (2024)
CausalX: Causal Explanations and Block Multilinear Factor Analysis
by: Vasilescu, M. Alex O., et al.
Published: (2021)
by: Vasilescu, M. Alex O., et al.
Published: (2021)
EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation
by: Li, Zilin, et al.
Published: (2026)
by: Li, Zilin, et al.
Published: (2026)
Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training
by: Gu, Zhengyao, et al.
Published: (2026)
by: Gu, Zhengyao, et al.
Published: (2026)
GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies
by: Zhang, He, et al.
Published: (2026)
by: Zhang, He, et al.
Published: (2026)
Similar Items
-
Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-Critic
by: Ji, Tianying, et al.
Published: (2023) -
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
by: Luo, Yu, et al.
Published: (2024) -
OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
by: Luo, Yu, et al.
Published: (2024) -
Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted Diffusion
by: Liang, Yongyuan, et al.
Published: (2024) -
Soft Actor-Critic with Beta Policy via Implicit Reparameterization Gradients
by: Della Libera, Luca
Published: (2024)