Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities
Fuente:
arXiv
Saved in:
| Main Authors: | Abraham, Armaan A., Shi, Lucy Xiaoyang, Finn, Chelsea |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
by: Guo, Yanjiang, et al.
Published: (2025)
by: Guo, Yanjiang, et al.
Published: (2025)
Learning Long-Context Diffusion Policies via Past-Token Prediction
by: Torne, Marcel, et al.
Published: (2025)
by: Torne, Marcel, et al.
Published: (2025)
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
by: Putta, Pranav, et al.
Published: (2024)
by: Putta, Pranav, et al.
Published: (2024)
Reinforcement Learning via Implicit Imitation Guidance
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
FASTER: Value-Guided Sampling for Fast RL
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
Latent Diffusion Planning for Imitation Learning
by: Xie, Amber, et al.
Published: (2025)
by: Xie, Amber, et al.
Published: (2025)
Value Flows
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
What Matters for Batch Online Reinforcement Learning in Robotics?
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
EXPO: Stable Reinforcement Learning with Expressive Policies
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Yell At Your Robot: Improving On-the-Fly from Language Corrections
by: Shi, Lucy Xiaoyang, et al.
Published: (2024)
by: Shi, Lucy Xiaoyang, et al.
Published: (2024)
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
by: Hsu, Sheryl, et al.
Published: (2024)
by: Hsu, Sheryl, et al.
Published: (2024)
Affordance-Guided Reinforcement Learning via Visual Prompting
by: Lee, Olivia Y., et al.
Published: (2024)
by: Lee, Olivia Y., et al.
Published: (2024)
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning
by: Xie, Johnathan, et al.
Published: (2024)
by: Xie, Johnathan, et al.
Published: (2024)
Prioritized Soft Q-Decomposition for Lexicographic Reinforcement Learning
by: Rietz, Finn, et al.
Published: (2023)
by: Rietz, Finn, et al.
Published: (2023)
Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
Learning Long-Horizon Predictions for Quadrotor Dynamics
by: Rao, Pratyaksh Prabhav, et al.
Published: (2024)
by: Rao, Pratyaksh Prabhav, et al.
Published: (2024)
Beyond Single-Step Updates: Reinforcement Learning of Heuristics with Limited-Horizon Search
by: Hadar, Gal, et al.
Published: (2025)
by: Hadar, Gal, et al.
Published: (2025)
Learning for Long-Horizon Planning via Neuro-Symbolic Abductive Imitation
by: Shao, Jie-Jing, et al.
Published: (2024)
by: Shao, Jie-Jing, et al.
Published: (2024)
Polychromic Objectives for Reinforcement Learning
by: Hamid, Jubayer Ibn, et al.
Published: (2025)
by: Hamid, Jubayer Ibn, et al.
Published: (2025)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
by: Wang, Zeyuan, et al.
Published: (2025)
by: Wang, Zeyuan, et al.
Published: (2025)
EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
Self-Guided Action Diffusion
by: Malhotra, Rhea, et al.
Published: (2025)
by: Malhotra, Rhea, et al.
Published: (2025)
Learning Agent-Compatible Context Management for Long-Horizon Tasks
by: Yi, Lu, et al.
Published: (2026)
by: Yi, Lu, et al.
Published: (2026)
DETACH: Cross-domain Learning for Long-Horizon Tasks via Mixture of Disentangled Experts
by: Shen, Yutong, et al.
Published: (2025)
by: Shen, Yutong, et al.
Published: (2025)
Universal Neural Functionals
by: Zhou, Allan, et al.
Published: (2024)
by: Zhou, Allan, et al.
Published: (2024)
Reinforcement Learning for Long-Horizon Interactive LLM Agents
by: Chen, Kevin, et al.
Published: (2025)
by: Chen, Kevin, et al.
Published: (2025)
MemER: Scaling Up Memory for Robot Control via Experience Retrieval
by: Sridhar, Ajay, et al.
Published: (2025)
by: Sridhar, Ajay, et al.
Published: (2025)
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)
by: Hejna, Joey, et al.
Published: (2023)
Toward Accurate Long-Horizon Robotic Manipulation: Language-to-Action with Foundation Models via Scene Graphs
by: Dinesh, Sushil Samuel, et al.
Published: (2025)
by: Dinesh, Sushil Samuel, et al.
Published: (2025)
Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
by: Fu, Zipeng, et al.
Published: (2024)
by: Fu, Zipeng, et al.
Published: (2024)
Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
by: Yang, Cheng, et al.
Published: (2025)
by: Yang, Cheng, et al.
Published: (2025)
Learning Bilevel Policies over Symbolic World Models for Long-Horizon Planning
by: Chen, Dillon Z., et al.
Published: (2026)
by: Chen, Dillon Z., et al.
Published: (2026)
GHQ: Grouped Hybrid Q Learning for Heterogeneous Cooperative Multi-agent Reinforcement Learning
by: Yu, Xiaoyang, et al.
Published: (2023)
by: Yu, Xiaoyang, et al.
Published: (2023)
Long-Horizon Visual Imitation Learning via Plan and Code Reflection
by: Chen, Quan, et al.
Published: (2025)
by: Chen, Quan, et al.
Published: (2025)
Milestone-Guided Policy Learning for Long-Horizon Language Agents
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
Alignment in Time: Peak-Aware Orchestration for Long-Horizon Agentic Systems
by: Shi, Hanjing, et al.
Published: (2026)
by: Shi, Hanjing, et al.
Published: (2026)
TSUBASA: Improving Long-Horizon Personalization via Evolving Memory and Self-Learning with Context Distillation
by: Zhang, Xinliang Frederick, et al.
Published: (2026)
by: Zhang, Xinliang Frederick, et al.
Published: (2026)
Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering
by: Zhu, Xinyu, et al.
Published: (2026)
by: Zhu, Xinyu, et al.
Published: (2026)
Reinforcement Learning for Long-Horizon Unordered Tasks: From Boolean to Coupled Reward Machines
by: Levina, Kristina, et al.
Published: (2025)
by: Levina, Kristina, et al.
Published: (2025)
Similar Items
-
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
by: Guo, Yanjiang, et al.
Published: (2025) -
Learning Long-Context Diffusion Policies via Past-Token Prediction
by: Torne, Marcel, et al.
Published: (2025) -
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
by: Putta, Pranav, et al.
Published: (2024) -
Reinforcement Learning via Implicit Imitation Guidance
by: Dong, Perry, et al.
Published: (2025) -
FASTER: Value-Guided Sampling for Fast RL
by: Dong, Perry, et al.
Published: (2026)