Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yao, Qingmao, Lei, Zhichao, Chen, Tianyuan, Yuan, Ziyue, Chen, Xuefan, Liu, Jianxiang, Wu, Faguo, Zhang, Xiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Diverse Policies with Soft Self-Generated Guidance
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression
von: Mao, Yixiu, et al.
Veröffentlicht: (2024)
von: Mao, Yixiu, et al.
Veröffentlicht: (2024)
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023)
von: Liu, Yao, et al.
Veröffentlicht: (2023)
A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs
von: Wei, Jiacheng, et al.
Veröffentlicht: (2025)
von: Wei, Jiacheng, et al.
Veröffentlicht: (2025)
Policy Optimization with Smooth Guidance Learned from State-Only Demonstrations
von: Wang, Guojian, et al.
Veröffentlicht: (2023)
von: Wang, Guojian, et al.
Veröffentlicht: (2023)
GAS: Enhancing Reward-Cost Balance of Generative Model-assisted Offline Safe RL
von: Liu, Zifan, et al.
Veröffentlicht: (2026)
von: Liu, Zifan, et al.
Veröffentlicht: (2026)
Towards Compositional Generalization in LLMs for Smart Contract Security: A Case Study on Reentrancy Vulnerabilities
von: Zhou, Ying, et al.
Veröffentlicht: (2026)
von: Zhou, Ying, et al.
Veröffentlicht: (2026)
Variational OOD State Correction for Offline Reinforcement Learning
von: Jiang, Ke, et al.
Veröffentlicht: (2025)
von: Jiang, Ke, et al.
Veröffentlicht: (2025)
RL Fine-Tuning Heals OOD Forgetting in SFT
von: Jin, Hangzhan, et al.
Veröffentlicht: (2025)
von: Jin, Hangzhan, et al.
Veröffentlicht: (2025)
Decoupled Prioritized Resampling for Offline RL
von: Yue, Yang, et al.
Veröffentlicht: (2023)
von: Yue, Yang, et al.
Veröffentlicht: (2023)
METRO: Towards Strategy Induction from Expert Dialogue Transcripts for Non-collaborative Dialogues
von: Yang, Haofu, et al.
Veröffentlicht: (2026)
von: Yang, Haofu, et al.
Veröffentlicht: (2026)
A Convex Hull Cheapest Insertion Heuristic for the Non-Euclidean TSP
von: Goutham, Mithun, et al.
Veröffentlicht: (2023)
von: Goutham, Mithun, et al.
Veröffentlicht: (2023)
Uncertainty Quantification in Large Language Models Through Convex Hull Analysis
von: Catak, Ferhat Ozgur, et al.
Veröffentlicht: (2024)
von: Catak, Ferhat Ozgur, et al.
Veröffentlicht: (2024)
A Tractable Inference Perspective of Offline RL
von: Liu, Xuejie, et al.
Veröffentlicht: (2023)
von: Liu, Xuejie, et al.
Veröffentlicht: (2023)
Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language Supervision
von: Zhang, Shilin, et al.
Veröffentlicht: (2025)
von: Zhang, Shilin, et al.
Veröffentlicht: (2025)
Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL
von: Zhong, Fangwei, et al.
Veröffentlicht: (2024)
von: Zhong, Fangwei, et al.
Veröffentlicht: (2024)
Adaptive Neighborhood-Constrained Q Learning for Offline Reinforcement Learning
von: Mao, Yixiu, et al.
Veröffentlicht: (2025)
von: Mao, Yixiu, et al.
Veröffentlicht: (2025)
Selective Uncertainty Propagation in Offline RL
von: Krishnamurthy, Sanath Kumar, et al.
Veröffentlicht: (2023)
von: Krishnamurthy, Sanath Kumar, et al.
Veröffentlicht: (2023)
Augmenting Offline RL with Unlabeled Data
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
von: Chen, Ziru, et al.
Veröffentlicht: (2026)
von: Chen, Ziru, et al.
Veröffentlicht: (2026)
Probabilistic Verification of Neural Networks via Efficient Probabilistic Hull Generation
von: Li, Jingyang, et al.
Veröffentlicht: (2026)
von: Li, Jingyang, et al.
Veröffentlicht: (2026)
Proto-OOD: Enhancing OOD Object Detection with Prototype Feature Similarity
von: Chen, Junkun, et al.
Veröffentlicht: (2024)
von: Chen, Junkun, et al.
Veröffentlicht: (2024)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
von: Su, Jianhai, et al.
Veröffentlicht: (2025)
von: Su, Jianhai, et al.
Veröffentlicht: (2025)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
von: Xiao, Wei, et al.
Veröffentlicht: (2025)
von: Xiao, Wei, et al.
Veröffentlicht: (2025)
STO-RL: Offline RL under Sparse Rewards via LLM-Guided Subgoal Temporal Order
von: Gu, Chengyang, et al.
Veröffentlicht: (2026)
von: Gu, Chengyang, et al.
Veröffentlicht: (2026)
Uncertainty Measurement of Deep Learning System based on the Convex Hull of Training Sets
von: Hwang, Hyekyoung, et al.
Veröffentlicht: (2024)
von: Hwang, Hyekyoung, et al.
Veröffentlicht: (2024)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
Design Considerations in Offline Preference-based RL
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
Are Expressive Models Truly Necessary for Offline RL?
von: Wang, Guan, et al.
Veröffentlicht: (2024)
von: Wang, Guan, et al.
Veröffentlicht: (2024)
OGBench: Benchmarking Offline Goal-Conditioned RL
von: Park, Seohong, et al.
Veröffentlicht: (2024)
von: Park, Seohong, et al.
Veröffentlicht: (2024)
Preference-Guided Reinforcement Learning for Efficient Exploration
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
Disentangling Policy from Offline Task Representation Learning via Adversarial Data Augmentation
von: Jia, Chengxing, et al.
Veröffentlicht: (2024)
von: Jia, Chengxing, et al.
Veröffentlicht: (2024)
Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
Improving Medical Diagnostics with Vision-Language Models: Convex Hull-Based Uncertainty Analysis
von: Catak, Ferhat Ozgur, et al.
Veröffentlicht: (2024)
von: Catak, Ferhat Ozgur, et al.
Veröffentlicht: (2024)
Improving Zero-Shot Offline RL via Behavioral Task Sampling
von: Bendib, Nazim, et al.
Veröffentlicht: (2026)
von: Bendib, Nazim, et al.
Veröffentlicht: (2026)
Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL
von: Dayal, Sarthak, et al.
Veröffentlicht: (2026)
von: Dayal, Sarthak, et al.
Veröffentlicht: (2026)
Offline Multi-task Transfer RL with Representational Penalization
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
Yes, Q-learning Helps Offline In-Context RL
von: Tarasov, Denis, et al.
Veröffentlicht: (2025)
von: Tarasov, Denis, et al.
Veröffentlicht: (2025)
Is Value Learning Really the Main Bottleneck in Offline RL?
von: Park, Seohong, et al.
Veröffentlicht: (2024)
von: Park, Seohong, et al.
Veröffentlicht: (2024)
The Role of Deep Learning Regularizations on Actors in Offline RL
von: Tarasov, Denis, et al.
Veröffentlicht: (2024)
von: Tarasov, Denis, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning Diverse Policies with Soft Self-Generated Guidance
von: Wang, Guojian, et al.
Veröffentlicht: (2024) -
Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression
von: Mao, Yixiu, et al.
Veröffentlicht: (2024) -
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023) -
A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs
von: Wei, Jiacheng, et al.
Veröffentlicht: (2025) -
Policy Optimization with Smooth Guidance Learned from State-Only Demonstrations
von: Wang, Guojian, et al.
Veröffentlicht: (2023)