Gespeichert in:
| Hauptverfasser: | Wang, Guojian, Wu, Faguo, Zhang, Xiao, Liu, Jianxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2402.04539 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Policy Optimization with Smooth Guidance Learned from State-Only Demonstrations
von: Wang, Guojian, et al.
Veröffentlicht: (2023)
von: Wang, Guojian, et al.
Veröffentlicht: (2023)
Trajectory-Oriented Policy Optimization with Sparse Rewards
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
von: Yao, Qingmao, et al.
Veröffentlicht: (2025)
von: Yao, Qingmao, et al.
Veröffentlicht: (2025)
Preference-Guided Reinforcement Learning for Efficient Exploration
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
Adaptive trajectory-constrained exploration strategy for deep reinforcement learning
von: Wang, Guojian, et al.
Veröffentlicht: (2023)
von: Wang, Guojian, et al.
Veröffentlicht: (2023)
Data Fusion-Enhanced Decision Transformer for Stable Cross-Domain Generalization
von: Wang, Guojian, et al.
Veröffentlicht: (2025)
von: Wang, Guojian, et al.
Veröffentlicht: (2025)
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
von: Zhan, Guojian, et al.
Veröffentlicht: (2026)
von: Zhan, Guojian, et al.
Veröffentlicht: (2026)
FFHFlow: Diverse and Uncertainty-Aware Dexterous Grasp Generation via Flow Variational Inference
von: Feng, Qian, et al.
Veröffentlicht: (2024)
von: Feng, Qian, et al.
Veröffentlicht: (2024)
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
von: Zhang, Feihong, et al.
Veröffentlicht: (2025)
von: Zhang, Feihong, et al.
Veröffentlicht: (2025)
VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning
von: Chen, Xuyang, et al.
Veröffentlicht: (2025)
von: Chen, Xuyang, et al.
Veröffentlicht: (2025)
Learning to Reason under Off-Policy Guidance
von: Yan, Jianhao, et al.
Veröffentlicht: (2025)
von: Yan, Jianhao, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Curriculum-inspired Adaptive Direct Policy Guidance for Truck Dispatching
von: Meng, Shi, et al.
Veröffentlicht: (2025)
von: Meng, Shi, et al.
Veröffentlicht: (2025)
Distributional Soft Actor-Critic with Diffusion Policy
von: Liu, Tong, et al.
Veröffentlicht: (2025)
von: Liu, Tong, et al.
Veröffentlicht: (2025)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
Self-Pro: A Self-Prompt and Tuning Framework for Graph Neural Networks
von: Gong, Chenghua, et al.
Veröffentlicht: (2023)
von: Gong, Chenghua, et al.
Veröffentlicht: (2023)
Frequency-Forcing: From Scaling-as-Time to Soft Frequency Guidance
von: Du, Weitao
Veröffentlicht: (2026)
von: Du, Weitao
Veröffentlicht: (2026)
Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance
von: He, Jinmin, et al.
Veröffentlicht: (2025)
von: He, Jinmin, et al.
Veröffentlicht: (2025)
More Than One Teacher: Adaptive Multi-Guidance Policy Optimization for Diverse Exploration
von: Yuan, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Yuan, Xiaoyang, et al.
Veröffentlicht: (2025)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
Soft Sequence Policy Optimization
von: Glazyrina, Svetlana, et al.
Veröffentlicht: (2026)
von: Glazyrina, Svetlana, et al.
Veröffentlicht: (2026)
Learning Policy Committees for Effective Personalization in MDPs with Diverse Tasks
von: Ge, Luise, et al.
Veröffentlicht: (2025)
von: Ge, Luise, et al.
Veröffentlicht: (2025)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance
von: He, Yufei, et al.
Veröffentlicht: (2025)
von: He, Yufei, et al.
Veröffentlicht: (2025)
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
von: Gao, Chen-Xiao, et al.
Veröffentlicht: (2025)
von: Gao, Chen-Xiao, et al.
Veröffentlicht: (2025)
Causal Policy Learning in Reinforcement Learning: Backdoor-Adjusted Soft Actor-Critic
von: Vo, Thanh Vinh, et al.
Veröffentlicht: (2025)
von: Vo, Thanh Vinh, et al.
Veröffentlicht: (2025)
Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following
von: Zeng, Yirong, et al.
Veröffentlicht: (2026)
von: Zeng, Yirong, et al.
Veröffentlicht: (2026)
FedRD: Reducing Divergences for Generalized Federated Learning via Heterogeneity-aware Parameter Guidance
von: Wang, Kaile, et al.
Veröffentlicht: (2026)
von: Wang, Kaile, et al.
Veröffentlicht: (2026)
A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs
von: Wei, Jiacheng, et al.
Veröffentlicht: (2025)
von: Wei, Jiacheng, et al.
Veröffentlicht: (2025)
AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
von: Liu, Ziru, et al.
Veröffentlicht: (2025)
von: Liu, Ziru, et al.
Veröffentlicht: (2025)
Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
von: Yang, Hanlin, et al.
Veröffentlicht: (2024)
von: Yang, Hanlin, et al.
Veröffentlicht: (2024)
Soft Adaptive Policy Optimization
von: Gao, Chang, et al.
Veröffentlicht: (2025)
von: Gao, Chang, et al.
Veröffentlicht: (2025)
Adaptive Guidance for Local Training in Heterogeneous Federated Learning
von: Zhang, Jianqing, et al.
Veröffentlicht: (2024)
von: Zhang, Jianqing, et al.
Veröffentlicht: (2024)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement Learning
von: Batra, Sumeet, et al.
Veröffentlicht: (2023)
von: Batra, Sumeet, et al.
Veröffentlicht: (2023)
Soft Deterministic Policy Gradient with Gaussian Smoothing
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
von: Wu, Yudi, et al.
Veröffentlicht: (2025)
von: Wu, Yudi, et al.
Veröffentlicht: (2025)
LRT-Diffusion: Calibrated Risk-Aware Guidance for Diffusion Policies
von: Sun, Ximan, et al.
Veröffentlicht: (2025)
von: Sun, Ximan, et al.
Veröffentlicht: (2025)
GTA: Generative Trajectory Augmentation with Guidance for Offline Reinforcement Learning
von: Lee, Jaewoo, et al.
Veröffentlicht: (2024)
von: Lee, Jaewoo, et al.
Veröffentlicht: (2024)
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
von: Reizinger, Patrik, et al.
Veröffentlicht: (2025)
von: Reizinger, Patrik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Policy Optimization with Smooth Guidance Learned from State-Only Demonstrations
von: Wang, Guojian, et al.
Veröffentlicht: (2023) -
Trajectory-Oriented Policy Optimization with Sparse Rewards
von: Wang, Guojian, et al.
Veröffentlicht: (2024) -
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
von: Yao, Qingmao, et al.
Veröffentlicht: (2025) -
Preference-Guided Reinforcement Learning for Efficient Exploration
von: Wang, Guojian, et al.
Veröffentlicht: (2024) -
Adaptive trajectory-constrained exploration strategy for deep reinforcement learning
von: Wang, Guojian, et al.
Veröffentlicht: (2023)