Debiased Model-based Representations for Sample-efficient Continuous Control
Fuente:
arXiv
Guardado en:
| Autores principales: | Lyu, Jiafei, Lin, Zichuan, Fujimoto, Scott, Yang, Kai, Chen, Yangkun, Yang, Saiyong, Lu, Zongqing, Ye, Deheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
por: Yang, Kai, et al.
Publicado: (2025)
por: Yang, Kai, et al.
Publicado: (2025)
Cross-Domain Policy Adaptation by Capturing Representation Mismatch
por: Lyu, Jiafei, et al.
Publicado: (2024)
por: Lyu, Jiafei, et al.
Publicado: (2024)
HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation
por: Zhao, He, et al.
Publicado: (2026)
por: Zhao, He, et al.
Publicado: (2026)
Understanding What Affects the Generalization Gap in Visual Reinforcement Learning: Theory and Empirical Evidence
por: Lyu, Jiafei, et al.
Publicado: (2024)
por: Lyu, Jiafei, et al.
Publicado: (2024)
Mildly Conservative Q-Learning for Offline Reinforcement Learning
por: Lyu, Jiafei, et al.
Publicado: (2022)
por: Lyu, Jiafei, et al.
Publicado: (2022)
ProAct: Agentic Lookahead in Interactive Environments
por: Yu, Yangbin, et al.
Publicado: (2026)
por: Yu, Yangbin, et al.
Publicado: (2026)
AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
por: Lin, Zichuan, et al.
Publicado: (2025)
por: Lin, Zichuan, et al.
Publicado: (2025)
GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training
por: Wei, Tong, et al.
Publicado: (2025)
por: Wei, Tong, et al.
Publicado: (2025)
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
por: Lin, Zichuan, et al.
Publicado: (2026)
por: Lin, Zichuan, et al.
Publicado: (2026)
Learning Versatile Skills with Curriculum Masking
por: Tang, Yao, et al.
Publicado: (2024)
por: Tang, Yao, et al.
Publicado: (2024)
PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning
por: Sun, Shengjie, et al.
Publicado: (2025)
por: Sun, Shengjie, et al.
Publicado: (2025)
A Two-stage Reinforcement Learning-based Approach for Multi-entity Task Allocation
por: Gong, Aicheng, et al.
Publicado: (2024)
por: Gong, Aicheng, et al.
Publicado: (2024)
CORD: Generalizable Cooperation via Role Diversity
por: Matsuyama, Kanefumi, et al.
Publicado: (2025)
por: Matsuyama, Kanefumi, et al.
Publicado: (2025)
CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasks
por: Chai, Qi, et al.
Publicado: (2025)
por: Chai, Qi, et al.
Publicado: (2025)
SEABO: A Simple Search-Based Method for Offline Imitation Learning
por: Lyu, Jiafei, et al.
Publicado: (2024)
por: Lyu, Jiafei, et al.
Publicado: (2024)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
por: Qu, Yun, et al.
Publicado: (2026)
por: Qu, Yun, et al.
Publicado: (2026)
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
por: Wei, Tong, et al.
Publicado: (2025)
por: Wei, Tong, et al.
Publicado: (2025)
Mind the Model, Not the Agent: The Primacy Bias in Model-based RL
por: Qiao, Zhongjian, et al.
Publicado: (2023)
por: Qiao, Zhongjian, et al.
Publicado: (2023)
PIPCFR: Pseudo-outcome Imputation with Post-treatment Variables for Individual Treatment Effect Estimation
por: Lin, Zichuan, et al.
Publicado: (2025)
por: Lin, Zichuan, et al.
Publicado: (2025)
ODRL: A Benchmark for Off-Dynamics Reinforcement Learning
por: Lyu, Jiafei, et al.
Publicado: (2024)
por: Lyu, Jiafei, et al.
Publicado: (2024)
Temporal Difference Learning with Constrained Initial Representations
por: Lyu, Jiafei, et al.
Publicado: (2026)
por: Lyu, Jiafei, et al.
Publicado: (2026)
The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
por: Yang, Zichuan
Publicado: (2025)
por: Yang, Zichuan
Publicado: (2025)
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
por: Qu, Yun, et al.
Publicado: (2026)
por: Qu, Yun, et al.
Publicado: (2026)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
por: Yang, Wenkai, et al.
Publicado: (2026)
por: Yang, Wenkai, et al.
Publicado: (2026)
Novelty-Guided Data Reuse for Efficient and Diversified Multi-Agent Reinforcement Learning
por: Chen, Yangkun, et al.
Publicado: (2024)
por: Chen, Yangkun, et al.
Publicado: (2024)
HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing
por: Chen, Ruyi, et al.
Publicado: (2026)
por: Chen, Ruyi, et al.
Publicado: (2026)
AdaRefiner: Refining Decisions of Language Models with Adaptive Feedback
por: Zhang, Wanpeng, et al.
Publicado: (2023)
por: Zhang, Wanpeng, et al.
Publicado: (2023)
WALL-E: World Alignment by Rule Learning Improves World Model-based LLM Agents
por: Zhou, Siyu, et al.
Publicado: (2024)
por: Zhou, Siyu, et al.
Publicado: (2024)
WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents
por: Zhou, Siyu, et al.
Publicado: (2025)
por: Zhou, Siyu, et al.
Publicado: (2025)
MTLight: Efficient Multi-Task Reinforcement Learning for Traffic Signal Control
por: Zhu, Liwen, et al.
Publicado: (2024)
por: Zhu, Liwen, et al.
Publicado: (2024)
Revisiting Discrete Soft Actor-Critic
por: Zhou, Haibin, et al.
Publicado: (2022)
por: Zhou, Haibin, et al.
Publicado: (2022)
Debiasing Reward Models by Representation Learning with Guarantees
por: Ng, Ignavier, et al.
Publicado: (2025)
por: Ng, Ignavier, et al.
Publicado: (2025)
SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
por: Wang, Yuxuan, et al.
Publicado: (2025)
por: Wang, Yuxuan, et al.
Publicado: (2025)
Vul-R2: A Reasoning LLM for Automated Vulnerability Repair
por: Wen, Xin-Cheng, et al.
Publicado: (2025)
por: Wen, Xin-Cheng, et al.
Publicado: (2025)
General Phrase Debiaser: Debiasing Masked Language Models at a Multi-Token Level
por: Shi, Bingkang, et al.
Publicado: (2023)
por: Shi, Bingkang, et al.
Publicado: (2023)
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
por: Lu, Zhicong, et al.
Publicado: (2026)
por: Lu, Zhicong, et al.
Publicado: (2026)
EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion
por: Yang, Zichuan, et al.
Publicado: (2025)
por: Yang, Zichuan, et al.
Publicado: (2025)
Causal Debiasing Medical Multimodal Representation Learning with Missing Modalities
por: Zhu, Xiaoguang, et al.
Publicado: (2025)
por: Zhu, Xiaoguang, et al.
Publicado: (2025)
Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary Dynamics
por: Zhang, Xinyu, et al.
Publicado: (2024)
por: Zhang, Xinyu, et al.
Publicado: (2024)
Boosting Vulnerability Detection of LLMs via Curriculum Preference Optimization with Synthetic Reasoning Data
por: Wen, Xin-Cheng, et al.
Publicado: (2025)
por: Wen, Xin-Cheng, et al.
Publicado: (2025)
Ejemplares similares
-
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
por: Yang, Kai, et al.
Publicado: (2025) -
Cross-Domain Policy Adaptation by Capturing Representation Mismatch
por: Lyu, Jiafei, et al.
Publicado: (2024) -
HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation
por: Zhao, He, et al.
Publicado: (2026) -
Understanding What Affects the Generalization Gap in Visual Reinforcement Learning: Theory and Empirical Evidence
por: Lyu, Jiafei, et al.
Publicado: (2024) -
Mildly Conservative Q-Learning for Offline Reinforcement Learning
por: Lyu, Jiafei, et al.
Publicado: (2022)