EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Chen, Wei, Lai, Zhang, Yanzhi, Shao, Chenyang, Dan, Zedong, Huang, Weiran, Zhang, Yuzhi, Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Targeted Exploration via Unified Entropy Control for Reinforcement Learning
by: Wang, Chen, et al.
Published: (2026)
by: Wang, Chen, et al.
Published: (2026)
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
Continual Offline Reinforcement Learning via Diffusion-based Dual Generative Replay
by: Liu, Jinmei, et al.
Published: (2024)
by: Liu, Jinmei, et al.
Published: (2024)
A Neural Model for Contextual Biasing Score Learning and Filtering
by: Huang, Wanting, et al.
Published: (2025)
by: Huang, Wanting, et al.
Published: (2025)
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
by: Xie, Can, et al.
Published: (2025)
by: Xie, Can, et al.
Published: (2025)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
Reasoning through Exploration: A Reinforcement Learning Framework for Robust Function Calling
by: Hao, Bingguang, et al.
Published: (2025)
by: Hao, Bingguang, et al.
Published: (2025)
Stateful Reasoning via Insight Replay
by: Lei, Bin, et al.
Published: (2026)
by: Lei, Bin, et al.
Published: (2026)
Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMs
by: Zhang, Wenjian, et al.
Published: (2026)
by: Zhang, Wenjian, et al.
Published: (2026)
DyJR: Preserving Diversity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay
by: Li, Long, et al.
Published: (2026)
by: Li, Long, et al.
Published: (2026)
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
by: Zhang, Kongcheng, et al.
Published: (2025)
by: Zhang, Kongcheng, et al.
Published: (2025)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
by: Deng, Wenhao, et al.
Published: (2025)
by: Deng, Wenhao, et al.
Published: (2025)
Internalizing LLM Reasoning via Discovery and Replay of Latent Actions
by: Shi, Zhenning, et al.
Published: (2026)
by: Shi, Zhenning, et al.
Published: (2026)
IDER: IDempotent Experience Replay for Reliable Continual Learning
by: Liu, Zhanwang, et al.
Published: (2026)
by: Liu, Zhanwang, et al.
Published: (2026)
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
by: Xu, Chengzhi, et al.
Published: (2025)
by: Xu, Chengzhi, et al.
Published: (2025)
Efficient Diversity-based Experience Replay for Deep Reinforcement Learning
by: Zhao, Kaiyan, et al.
Published: (2024)
by: Zhao, Kaiyan, et al.
Published: (2024)
Reinforced Efficient Reasoning via Semantically Diverse Exploration
by: Zhao, Ziqi, et al.
Published: (2026)
by: Zhao, Ziqi, et al.
Published: (2026)
Automated Design and Optimization of Distributed Filtering Circuits via Reinforcement Learning
by: Gao, Peng, et al.
Published: (2024)
by: Gao, Peng, et al.
Published: (2024)
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
by: Zhang, Yanzhi, et al.
Published: (2025)
by: Zhang, Yanzhi, et al.
Published: (2025)
Route-and-Reason: Scaling Large Language Model Reasoning with Reinforced Model Router
by: Shao, Chenyang, et al.
Published: (2025)
by: Shao, Chenyang, et al.
Published: (2025)
Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning
by: Liu, Jinyi, et al.
Published: (2023)
by: Liu, Jinyi, et al.
Published: (2023)
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
Deeper Insights into Learning Performance of Stochastic Configuration Networks
by: Yan, Xiufeng, et al.
Published: (2024)
by: Yan, Xiufeng, et al.
Published: (2024)
When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making
by: Liu, Jun, et al.
Published: (2026)
by: Liu, Jun, et al.
Published: (2026)
GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning
by: Chu, Xiangxiang, et al.
Published: (2025)
by: Chu, Xiangxiang, et al.
Published: (2025)
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
by: Wu, Zhiheng, et al.
Published: (2026)
by: Wu, Zhiheng, et al.
Published: (2026)
CIER: A Novel Experience Replay Approach with Causal Inference in Deep Reinforcement Learning
by: Wang, Jingwen, et al.
Published: (2024)
by: Wang, Jingwen, et al.
Published: (2024)
Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
TIMRL: A Novel Meta-Reinforcement Learning Framework for Non-Stationary and Multi-Task Environments
by: Qi, Chenyang, et al.
Published: (2025)
by: Qi, Chenyang, et al.
Published: (2025)
PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning
by: Guo, Weiran, et al.
Published: (2025)
by: Guo, Weiran, et al.
Published: (2025)
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization
by: Hua, Xingyuan, et al.
Published: (2026)
by: Hua, Xingyuan, et al.
Published: (2026)
FedReplay: A Feature Replay Assisted Federated Transfer Learning Framework for Efficient and Privacy-Preserving Smart Agriculture
by: Li, Long, et al.
Published: (2025)
by: Li, Long, et al.
Published: (2025)
Brain-Like Replay Naturally Emerges in Reinforcement Learning Agents
by: Wang, Jiyi, et al.
Published: (2024)
by: Wang, Jiyi, et al.
Published: (2024)
Guardian: Decoupling Exploration from Safety in Reinforcement Learning
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
MEGG: Replay via Maximally Extreme GGscore in Incremental Learning for Neural Recommendation Models
by: Shi, Yunxiao, et al.
Published: (2025)
by: Shi, Yunxiao, et al.
Published: (2025)
Credit Where It is Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning
by: Jiao, Zhengbo, et al.
Published: (2026)
by: Jiao, Zhengbo, et al.
Published: (2026)
Off-policy Reinforcement Learning with Model-based Exploration Augmentation
by: Wang, Likun, et al.
Published: (2025)
by: Wang, Likun, et al.
Published: (2025)
Experience Replay Addresses Loss of Plasticity in Continual Learning
by: Wang, Jiuqi, et al.
Published: (2025)
by: Wang, Jiuqi, et al.
Published: (2025)
Similar Items
-
Targeted Exploration via Unified Entropy Control for Reinforcement Learning
by: Wang, Chen, et al.
Published: (2026) -
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
by: Wei, Lai, et al.
Published: (2025) -
Continual Offline Reinforcement Learning via Diffusion-based Dual Generative Replay
by: Liu, Jinmei, et al.
Published: (2024) -
A Neural Model for Contextual Biasing Score Learning and Filtering
by: Huang, Wanting, et al.
Published: (2025) -
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
by: Xie, Can, et al.
Published: (2025)