Snapshot Reinforcement Learning: Leveraging Prior Trajectories for Efficiency
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Yanxiao, Qian, Yangge, Wang, Tianyi, Shan, Jingyang, Qin, Xiaolin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CAMEL: Continuous Action Masking Enabled by Large Language Models for Reinforcement Learning
por: Zhao, Yanxiao, et al.
Publicado: (2025)
por: Zhao, Yanxiao, et al.
Publicado: (2025)
SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
por: Zhao, Yanxiao, et al.
Publicado: (2025)
por: Zhao, Yanxiao, et al.
Publicado: (2025)
A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning
por: Hu, Yuzheng, et al.
Publicado: (2025)
por: Hu, Yuzheng, et al.
Publicado: (2025)
Criticality Leveraged Adversarial Training (CLAT) for Boosted Performance via Parameter Efficiency
por: Gopal, Bhavna, et al.
Publicado: (2024)
por: Gopal, Bhavna, et al.
Publicado: (2024)
A Model-Based Approach for Improving Reinforcement Learning Efficiency Leveraging Expert Observations
por: Ozcan, Erhan Can, et al.
Publicado: (2024)
por: Ozcan, Erhan Can, et al.
Publicado: (2024)
Enhancing Solution Efficiency in Reinforcement Learning: Leveraging Sub-GFlowNet and Entropy Integration
por: He, Siyi
Publicado: (2024)
por: He, Siyi
Publicado: (2024)
DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation
por: Xiang, Jingyang, et al.
Publicado: (2024)
por: Xiang, Jingyang, et al.
Publicado: (2024)
Efficient Reinforcement Learning with Large Language Model Priors
por: Yan, Xue, et al.
Publicado: (2024)
por: Yan, Xue, et al.
Publicado: (2024)
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching
por: Li, Guanghe, et al.
Publicado: (2024)
por: Li, Guanghe, et al.
Publicado: (2024)
Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation
por: Hu, Haichen, et al.
Publicado: (2026)
por: Hu, Haichen, et al.
Publicado: (2026)
In-Trajectory Inverse Reinforcement Learning: Learn Incrementally Before An Ongoing Trajectory Terminates
por: Liu, Shicheng, et al.
Publicado: (2024)
por: Liu, Shicheng, et al.
Publicado: (2024)
Anytime-Competitive Reinforcement Learning with Policy Prior
por: Yang, Jianyi, et al.
Publicado: (2023)
por: Yang, Jianyi, et al.
Publicado: (2023)
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
por: Liu, Wenpu, et al.
Publicado: (2026)
por: Liu, Wenpu, et al.
Publicado: (2026)
Learning from Snapshots of Discrete and Continuous Data Streams
por: Devulapalli, Pramith, et al.
Publicado: (2024)
por: Devulapalli, Pramith, et al.
Publicado: (2024)
Inverse Reinforcement Learning with Switching Rewards and History Dependency for Characterizing Animal Behaviors
por: Ke, Jingyang, et al.
Publicado: (2025)
por: Ke, Jingyang, et al.
Publicado: (2025)
Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement Learning
por: Ma, Hao, et al.
Publicado: (2025)
por: Ma, Hao, et al.
Publicado: (2025)
Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning?
por: Dai, Yang, et al.
Publicado: (2024)
por: Dai, Yang, et al.
Publicado: (2024)
Offline Trajectory Optimization for Offline Reinforcement Learning
por: Zhao, Ziqi, et al.
Publicado: (2024)
por: Zhao, Ziqi, et al.
Publicado: (2024)
Trajectory Entropy Reinforcement Learning for Predictable and Robust Control
por: You, Bang, et al.
Publicado: (2025)
por: You, Bang, et al.
Publicado: (2025)
Q-Pensieve: Boosting Sample Efficiency of Multi-Objective RL Through Memory Sharing of Q-Snapshots
por: Hung, Wei, et al.
Publicado: (2022)
por: Hung, Wei, et al.
Publicado: (2022)
Offline Reinforcement Learning with Generative Trajectory Policies
por: Feng, Xinsong, et al.
Publicado: (2025)
por: Feng, Xinsong, et al.
Publicado: (2025)
CellStream: Dynamical Optimal Transport Informed Embeddings for Reconstructing Cellular Trajectories from Snapshots Data
por: Ling, Yue, et al.
Publicado: (2025)
por: Ling, Yue, et al.
Publicado: (2025)
Efficient Online Reinforcement Learning for Diffusion Policy
por: Ma, Haitong, et al.
Publicado: (2025)
por: Ma, Haitong, et al.
Publicado: (2025)
Prior-Guided Diffusion Planning for Offline Reinforcement Learning
por: Ki, Donghyeon, et al.
Publicado: (2025)
por: Ki, Donghyeon, et al.
Publicado: (2025)
Query-Policy Misalignment in Preference-Based Reinforcement Learning
por: Hu, Xiao, et al.
Publicado: (2023)
por: Hu, Xiao, et al.
Publicado: (2023)
Policy-Based Trajectory Clustering in Offline Reinforcement Learning
por: Hu, Hao, et al.
Publicado: (2025)
por: Hu, Hao, et al.
Publicado: (2025)
PRCD-MAP: Learning How Much to Trust Imperfect Priors in Causal Discovery
por: Shan, Xihang, et al.
Publicado: (2026)
por: Shan, Xihang, et al.
Publicado: (2026)
Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement Learning
por: Mutti, Mirco, et al.
Publicado: (2023)
por: Mutti, Mirco, et al.
Publicado: (2023)
Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
por: Shi, Taiwei, et al.
Publicado: (2025)
por: Shi, Taiwei, et al.
Publicado: (2025)
Sparse Threats, Focused Defense: Criticality-Aware Robust Reinforcement Learning for Safe Autonomous Driving
por: Wei, Qi, et al.
Publicado: (2026)
por: Wei, Qi, et al.
Publicado: (2026)
ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
por: Zhang, Feng, et al.
Publicado: (2025)
por: Zhang, Feng, et al.
Publicado: (2025)
Enhancing Offline Reinforcement Learning with Curriculum Learning-Based Trajectory Valuation
por: Abolfazli, Amir, et al.
Publicado: (2025)
por: Abolfazli, Amir, et al.
Publicado: (2025)
Know your Trajectory -- Trustworthy Reinforcement Learning deployment through Importance-Based Trajectory Analysis
por: F, Clifford, et al.
Publicado: (2025)
por: F, Clifford, et al.
Publicado: (2025)
Enhancing Reinforcement Learning Fine-Tuning with an Online Refiner
por: Ma, Hao, et al.
Publicado: (2026)
por: Ma, Hao, et al.
Publicado: (2026)
Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement Learning
por: Daley, Brett, et al.
Publicado: (2023)
por: Daley, Brett, et al.
Publicado: (2023)
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning
por: Xu, Tianshi, et al.
Publicado: (2026)
por: Xu, Tianshi, et al.
Publicado: (2026)
Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning
por: Beylier, Charlotte, et al.
Publicado: (2025)
por: Beylier, Charlotte, et al.
Publicado: (2025)
A Multimodal Cross-View Model for Predicting Postoperative Neck Pain in Cervical Spondylosis Patients
por: Shan, Jingyang, et al.
Publicado: (2025)
por: Shan, Jingyang, et al.
Publicado: (2025)
ReFill: Reinforcement Learning for Fill-In Minimization
por: Harb, Elfarouk, et al.
Publicado: (2025)
por: Harb, Elfarouk, et al.
Publicado: (2025)
Orthogonal Weight Modification Enhances Learning Scalability and Convergence Efficiency without Gradient Backpropagation
por: Ma, Guoqing, et al.
Publicado: (2026)
por: Ma, Guoqing, et al.
Publicado: (2026)
Ejemplares similares
-
CAMEL: Continuous Action Masking Enabled by Large Language Models for Reinforcement Learning
por: Zhao, Yanxiao, et al.
Publicado: (2025) -
SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
por: Zhao, Yanxiao, et al.
Publicado: (2025) -
A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning
por: Hu, Yuzheng, et al.
Publicado: (2025) -
Criticality Leveraged Adversarial Training (CLAT) for Boosted Performance via Parameter Efficiency
por: Gopal, Bhavna, et al.
Publicado: (2024) -
A Model-Based Approach for Improving Reinforcement Learning Efficiency Leveraging Expert Observations
por: Ozcan, Erhan Can, et al.
Publicado: (2024)