Debiased Model-based Representations for Sample-efficient Continuous Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lyu, Jiafei, Lin, Zichuan, Fujimoto, Scott, Yang, Kai, Chen, Yangkun, Yang, Saiyong, Lu, Zongqing, Ye, Deheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
von: Yang, Kai, et al.
Veröffentlicht: (2025)
von: Yang, Kai, et al.
Veröffentlicht: (2025)
Cross-Domain Policy Adaptation by Capturing Representation Mismatch
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation
von: Zhao, He, et al.
Veröffentlicht: (2026)
von: Zhao, He, et al.
Veröffentlicht: (2026)
Understanding What Affects the Generalization Gap in Visual Reinforcement Learning: Theory and Empirical Evidence
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
Mildly Conservative Q-Learning for Offline Reinforcement Learning
von: Lyu, Jiafei, et al.
Veröffentlicht: (2022)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2022)
ProAct: Agentic Lookahead in Interactive Environments
von: Yu, Yangbin, et al.
Veröffentlicht: (2026)
von: Yu, Yangbin, et al.
Veröffentlicht: (2026)
AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
von: Lin, Zichuan, et al.
Veröffentlicht: (2025)
von: Lin, Zichuan, et al.
Veröffentlicht: (2025)
GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training
von: Wei, Tong, et al.
Veröffentlicht: (2025)
von: Wei, Tong, et al.
Veröffentlicht: (2025)
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
von: Lin, Zichuan, et al.
Veröffentlicht: (2026)
von: Lin, Zichuan, et al.
Veröffentlicht: (2026)
Learning Versatile Skills with Curriculum Masking
von: Tang, Yao, et al.
Veröffentlicht: (2024)
von: Tang, Yao, et al.
Veröffentlicht: (2024)
PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning
von: Sun, Shengjie, et al.
Veröffentlicht: (2025)
von: Sun, Shengjie, et al.
Veröffentlicht: (2025)
A Two-stage Reinforcement Learning-based Approach for Multi-entity Task Allocation
von: Gong, Aicheng, et al.
Veröffentlicht: (2024)
von: Gong, Aicheng, et al.
Veröffentlicht: (2024)
CORD: Generalizable Cooperation via Role Diversity
von: Matsuyama, Kanefumi, et al.
Veröffentlicht: (2025)
von: Matsuyama, Kanefumi, et al.
Veröffentlicht: (2025)
CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasks
von: Chai, Qi, et al.
Veröffentlicht: (2025)
von: Chai, Qi, et al.
Veröffentlicht: (2025)
SEABO: A Simple Search-Based Method for Offline Imitation Learning
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
von: Qu, Yun, et al.
Veröffentlicht: (2026)
von: Qu, Yun, et al.
Veröffentlicht: (2026)
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
von: Wei, Tong, et al.
Veröffentlicht: (2025)
von: Wei, Tong, et al.
Veröffentlicht: (2025)
Mind the Model, Not the Agent: The Primacy Bias in Model-based RL
von: Qiao, Zhongjian, et al.
Veröffentlicht: (2023)
von: Qiao, Zhongjian, et al.
Veröffentlicht: (2023)
PIPCFR: Pseudo-outcome Imputation with Post-treatment Variables for Individual Treatment Effect Estimation
von: Lin, Zichuan, et al.
Veröffentlicht: (2025)
von: Lin, Zichuan, et al.
Veröffentlicht: (2025)
ODRL: A Benchmark for Off-Dynamics Reinforcement Learning
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
Temporal Difference Learning with Constrained Initial Representations
von: Lyu, Jiafei, et al.
Veröffentlicht: (2026)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2026)
The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
von: Yang, Zichuan
Veröffentlicht: (2025)
von: Yang, Zichuan
Veröffentlicht: (2025)
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
von: Qu, Yun, et al.
Veröffentlicht: (2026)
von: Qu, Yun, et al.
Veröffentlicht: (2026)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
Novelty-Guided Data Reuse for Efficient and Diversified Multi-Agent Reinforcement Learning
von: Chen, Yangkun, et al.
Veröffentlicht: (2024)
von: Chen, Yangkun, et al.
Veröffentlicht: (2024)
HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing
von: Chen, Ruyi, et al.
Veröffentlicht: (2026)
von: Chen, Ruyi, et al.
Veröffentlicht: (2026)
AdaRefiner: Refining Decisions of Language Models with Adaptive Feedback
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2023)
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2023)
WALL-E: World Alignment by Rule Learning Improves World Model-based LLM Agents
von: Zhou, Siyu, et al.
Veröffentlicht: (2024)
von: Zhou, Siyu, et al.
Veröffentlicht: (2024)
WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents
von: Zhou, Siyu, et al.
Veröffentlicht: (2025)
von: Zhou, Siyu, et al.
Veröffentlicht: (2025)
MTLight: Efficient Multi-Task Reinforcement Learning for Traffic Signal Control
von: Zhu, Liwen, et al.
Veröffentlicht: (2024)
von: Zhu, Liwen, et al.
Veröffentlicht: (2024)
Revisiting Discrete Soft Actor-Critic
von: Zhou, Haibin, et al.
Veröffentlicht: (2022)
von: Zhou, Haibin, et al.
Veröffentlicht: (2022)
Debiasing Reward Models by Representation Learning with Guarantees
von: Ng, Ignavier, et al.
Veröffentlicht: (2025)
von: Ng, Ignavier, et al.
Veröffentlicht: (2025)
SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
von: Wang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2025)
Vul-R2: A Reasoning LLM for Automated Vulnerability Repair
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2025)
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2025)
General Phrase Debiaser: Debiasing Masked Language Models at a Multi-Token Level
von: Shi, Bingkang, et al.
Veröffentlicht: (2023)
von: Shi, Bingkang, et al.
Veröffentlicht: (2023)
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
von: Lu, Zhicong, et al.
Veröffentlicht: (2026)
von: Lu, Zhicong, et al.
Veröffentlicht: (2026)
EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion
von: Yang, Zichuan, et al.
Veröffentlicht: (2025)
von: Yang, Zichuan, et al.
Veröffentlicht: (2025)
Causal Debiasing Medical Multimodal Representation Learning with Missing Modalities
von: Zhu, Xiaoguang, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoguang, et al.
Veröffentlicht: (2025)
Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary Dynamics
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
Boosting Vulnerability Detection of LLMs via Curriculum Preference Optimization with Synthetic Reasoning Data
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2025)
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
von: Yang, Kai, et al.
Veröffentlicht: (2025) -
Cross-Domain Policy Adaptation by Capturing Representation Mismatch
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024) -
HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation
von: Zhao, He, et al.
Veröffentlicht: (2026) -
Understanding What Affects the Generalization Gap in Visual Reinforcement Learning: Theory and Empirical Evidence
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024) -
Mildly Conservative Q-Learning for Offline Reinforcement Learning
von: Lyu, Jiafei, et al.
Veröffentlicht: (2022)