Poly-EPO: Training Exploratory Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Orney, Ifdita Hasan, Hamid, Jubayer Ibn, Ramanujam, Shreya S, Wu, Shirley, Hu, Hengyuan, Goodman, Noah, Sadigh, Dorsa, Finn, Chelsea |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Polychromic Objectives for Reinforcement Learning
by: Hamid, Jubayer Ibn, et al.
Published: (2025)
by: Hamid, Jubayer Ibn, et al.
Published: (2025)
Invariance Co-training for Robot Visual Generalization
by: Yang, Jonathan, et al.
Published: (2025)
by: Yang, Jonathan, et al.
Published: (2025)
Imitation Bootstrapped Reinforcement Learning
by: Hu, Hengyuan, et al.
Published: (2023)
by: Hu, Hengyuan, et al.
Published: (2023)
Latent Diffusion Planning for Imitation Learning
by: Xie, Amber, et al.
Published: (2025)
by: Xie, Amber, et al.
Published: (2025)
What Matters for Batch Online Reinforcement Learning in Robotics?
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling
by: Liu, Yuejiang, et al.
Published: (2024)
by: Liu, Yuejiang, et al.
Published: (2024)
EXPO: Stable Reinforcement Learning with Expressive Policies
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
FASTER: Value-Guided Sampling for Fast RL
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Autospeculation
by: Hu, Hengyuan, et al.
Published: (2025)
by: Hu, Hengyuan, et al.
Published: (2025)
Toward Grounded Commonsense Reasoning
by: Kwon, Minae, et al.
Published: (2023)
by: Kwon, Minae, et al.
Published: (2023)
Neural Garbage Collection: Learning to Forget while Learning to Reason
by: Li, Michael Y., et al.
Published: (2026)
by: Li, Michael Y., et al.
Published: (2026)
Tripod: Three Complementary Inductive Biases for Disentangled Representation Learning
by: Hsu, Kyle, et al.
Published: (2024)
by: Hsu, Kyle, et al.
Published: (2024)
EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
RoboCade: Gamifying Robot Data Collection
by: Mirchandani, Suvir, et al.
Published: (2025)
by: Mirchandani, Suvir, et al.
Published: (2025)
Value Flows
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Efficient Data Collection for Robotic Manipulation via Compositional Generalization
by: Gao, Jensen, et al.
Published: (2024)
by: Gao, Jensen, et al.
Published: (2024)
Data Analogies Enable Efficient Cross-Embodiment Transfer
by: Yang, Jonathan, et al.
Published: (2026)
by: Yang, Jonathan, et al.
Published: (2026)
TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
What's the Move? Hybrid Imitation Learning via Salient Points
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
Action-Free Reasoning for Policy Generalization
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)
by: Hejna, Joey, et al.
Published: (2023)
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
by: Sarkar, Bidipta, et al.
Published: (2025)
by: Sarkar, Bidipta, et al.
Published: (2025)
Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
Can LLM-Simulated Practice and Feedback Upskill Human Counselors? A Randomized Study with 90+ Novice Counselors
by: Louie, Ryan, et al.
Published: (2025)
by: Louie, Ryan, et al.
Published: (2025)
Large Language Model Reasoning Failures
by: Song, Peiyang, et al.
Published: (2026)
by: Song, Peiyang, et al.
Published: (2026)
Data Retrieval with Importance Weights for Few-Shot Imitation Learning
by: Xie, Amber, et al.
Published: (2025)
by: Xie, Amber, et al.
Published: (2025)
Policy Learning with a Language Bottleneck
by: Srivastava, Megha, et al.
Published: (2024)
by: Srivastava, Megha, et al.
Published: (2024)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
by: Mishra, Shubhra, et al.
Published: (2024)
by: Mishra, Shubhra, et al.
Published: (2024)
MotIF: Motion Instruction Fine-tuning
by: Hwang, Minyoung, et al.
Published: (2024)
by: Hwang, Minyoung, et al.
Published: (2024)
ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2025)
by: Grannen, Jennifer, et al.
Published: (2025)
How to Train Your Robots? The Impact of Demonstration Modality on Imitation Learning
by: Li, Haozhuo, et al.
Published: (2025)
by: Li, Haozhuo, et al.
Published: (2025)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
by: Qu, Yuxiao, et al.
Published: (2025)
by: Qu, Yuxiao, et al.
Published: (2025)
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
by: Gao, Tian, et al.
Published: (2026)
by: Gao, Tian, et al.
Published: (2026)
Is Child-Directed Speech Effective Training Data for Language Models?
by: Feng, Steven Y., et al.
Published: (2024)
by: Feng, Steven Y., et al.
Published: (2024)
Motion Tracks: A Unified Representation for Human-Robot Transfer in Few-Shot Imitation Learning
by: Ren, Juntao, et al.
Published: (2025)
by: Ren, Juntao, et al.
Published: (2025)
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
by: Zou, Chelsea, et al.
Published: (2026)
by: Zou, Chelsea, et al.
Published: (2026)
Altruistic Maneuver Planning for Cooperative Autonomous Vehicles Using Multi-agent Advantage Actor-Critic
by: Toghi, Behrad, et al.
Published: (2021)
by: Toghi, Behrad, et al.
Published: (2021)
FLAIR: Feeding via Long-horizon AcquIsition of Realistic dishes
by: Jenamani, Rajat Kumar, et al.
Published: (2024)
by: Jenamani, Rajat Kumar, et al.
Published: (2024)
GIANTS: Generative Insight Anticipation from Scientific Literature
by: He-Yueya, Joy, et al.
Published: (2026)
by: He-Yueya, Joy, et al.
Published: (2026)
Similar Items
-
Polychromic Objectives for Reinforcement Learning
by: Hamid, Jubayer Ibn, et al.
Published: (2025) -
Invariance Co-training for Robot Visual Generalization
by: Yang, Jonathan, et al.
Published: (2025) -
Imitation Bootstrapped Reinforcement Learning
by: Hu, Hengyuan, et al.
Published: (2023) -
Latent Diffusion Planning for Imitation Learning
by: Xie, Amber, et al.
Published: (2025) -
What Matters for Batch Online Reinforcement Learning in Robotics?
by: Dong, Perry, et al.
Published: (2025)