Exploration by Random Reward Perturbation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Haozhe, Fu, Guoji, Luo, Zhengding, Wu, Jiele, Leong, Tze-Yun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning
by: Ma, Haozhe, et al.
Published: (2024)
by: Ma, Haozhe, et al.
Published: (2024)
Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement Learning
by: Ma, Haozhe, et al.
Published: (2024)
by: Ma, Haozhe, et al.
Published: (2024)
Hierarchical Molecular Representation Learning via Fragment-Based Self-Supervised Embedding Prediction
by: Wu, Jiele, et al.
Published: (2026)
by: Wu, Jiele, et al.
Published: (2026)
Causal Policy Learning in Reinforcement Learning: Backdoor-Adjusted Soft Actor-Critic
by: Vo, Thanh Vinh, et al.
Published: (2025)
by: Vo, Thanh Vinh, et al.
Published: (2025)
Performance Asymmetry in Model-Based Reinforcement Learning
by: Lim, Jing Yu, et al.
Published: (2025)
by: Lim, Jing Yu, et al.
Published: (2025)
Federated Causal Inference from Observational Data
by: Vo, Thanh Vinh, et al.
Published: (2023)
by: Vo, Thanh Vinh, et al.
Published: (2023)
Provably Efficient Exploration in Reward Machines with Low Regret
by: Bourel, Hippolyte, et al.
Published: (2024)
by: Bourel, Hippolyte, et al.
Published: (2024)
FDRMFL:Multi-modal Federated Feature Extraction Model Based on Information Maximization and Contrastive Learning
by: Wu, Haozhe
Published: (2025)
by: Wu, Haozhe
Published: (2025)
Decoupled Prompt-Adapter Tuning for Continual Activity Recognition
by: Fu, Di, et al.
Published: (2024)
by: Fu, Di, et al.
Published: (2024)
Continual Reinforcement Learning by Planning with Online World Models
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
by: Wang, Haozhe, et al.
Published: (2026)
by: Wang, Haozhe, et al.
Published: (2026)
MIR: Efficient Exploration in Episodic Multi-Agent Reinforcement Learning via Mutual Intrinsic Reward
by: Chen, Kesheng, et al.
Published: (2025)
by: Chen, Kesheng, et al.
Published: (2025)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
by: Russo, Alessio, et al.
Published: (2025)
by: Russo, Alessio, et al.
Published: (2025)
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
by: Sana, Mohamed, et al.
Published: (2026)
by: Sana, Mohamed, et al.
Published: (2026)
Exploration Through Introspection: A Self-Aware Reward Model
by: Petrowski, Michael, et al.
Published: (2026)
by: Petrowski, Michael, et al.
Published: (2026)
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
by: Zhang, Junkai, et al.
Published: (2024)
by: Zhang, Junkai, et al.
Published: (2024)
TSSR: Two-Stage Swap-Reward-Driven Reinforcement Learning for Character-Level SMILES Generation
by: Levine, Jacob Ede, et al.
Published: (2026)
by: Levine, Jacob Ede, et al.
Published: (2026)
What Are Step-Level Reward Models Rewarding? Counterintuitive Findings from MCTS-Boosted Mathematical Reasoning
by: Ma, Yiran, et al.
Published: (2024)
by: Ma, Yiran, et al.
Published: (2024)
BaNEL: Exploration Posteriors for Generative Modeling Using Only Negative Rewards
by: Lee, Sangyun, et al.
Published: (2025)
by: Lee, Sangyun, et al.
Published: (2025)
Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic Rewards
by: Mohamed, Faisal, et al.
Published: (2026)
by: Mohamed, Faisal, et al.
Published: (2026)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
by: Wang, Yangyue, et al.
Published: (2026)
by: Wang, Yangyue, et al.
Published: (2026)
Nonsense Helps: Prompt Space Perturbation Broadens Reasoning Exploration
by: Huang, Langlin, et al.
Published: (2026)
by: Huang, Langlin, et al.
Published: (2026)
Robot Policy Learning with Temporal Optimal Transport Reward
by: Fu, Yuwei, et al.
Published: (2024)
by: Fu, Yuwei, et al.
Published: (2024)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
by: Patel, Bhrij, et al.
Published: (2023)
by: Patel, Bhrij, et al.
Published: (2023)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
by: Ma, Rachel, et al.
Published: (2026)
by: Ma, Rachel, et al.
Published: (2026)
APEX: Probing Neural Networks via Activation Perturbation
by: Ren, Tao, et al.
Published: (2026)
by: Ren, Tao, et al.
Published: (2026)
More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling
by: Ishfaq, Haque, et al.
Published: (2024)
by: Ishfaq, Haque, et al.
Published: (2024)
MemReward: Graph-Based Experience Memory for LLM Reward Prediction with Limited Labels
by: Luo, Tianyang, et al.
Published: (2026)
by: Luo, Tianyang, et al.
Published: (2026)
SemiReward: A General Reward Model for Semi-supervised Learning
by: Li, Siyuan, et al.
Published: (2023)
by: Li, Siyuan, et al.
Published: (2023)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
Explaining Time Series via Contrastive and Locally Sparse Perturbations
by: Liu, Zichuan, et al.
Published: (2024)
by: Liu, Zichuan, et al.
Published: (2024)
Robust Offline Reinforcement learning with Heavy-Tailed Rewards
by: Zhu, Jin, et al.
Published: (2023)
by: Zhu, Jin, et al.
Published: (2023)
A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals
by: Liu, Grace, et al.
Published: (2024)
by: Liu, Grace, et al.
Published: (2024)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Unifying Adversarial Perturbation for Graph Neural Networks
by: Yang, Jinluan, et al.
Published: (2025)
by: Yang, Jinluan, et al.
Published: (2025)
MAESTRO: Multi-Agent Environment Shaping through Task and Reward Optimization
by: Wu, Boyuan
Published: (2025)
by: Wu, Boyuan
Published: (2025)
Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning
by: Qu, Yun, et al.
Published: (2024)
by: Qu, Yun, et al.
Published: (2024)
Similar Items
-
Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning
by: Ma, Haozhe, et al.
Published: (2024) -
Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement Learning
by: Ma, Haozhe, et al.
Published: (2024) -
Hierarchical Molecular Representation Learning via Fragment-Based Self-Supervised Embedding Prediction
by: Wu, Jiele, et al.
Published: (2026) -
Causal Policy Learning in Reinforcement Learning: Backdoor-Adjusted Soft Actor-Critic
by: Vo, Thanh Vinh, et al.
Published: (2025) -
Performance Asymmetry in Model-Based Reinforcement Learning
by: Lim, Jing Yu, et al.
Published: (2025)