Scaling Offline RL via Efficient and Expressive Shortcut Models
Fuente:
arXiv
Saved in:
| Main Authors: | Espinosa-Dice, Nicolas, Zhang, Yiyi, Chen, Yiding, Guo, Bradley, Oertell, Owen, Swamy, Gokul, Brantley, Kiante, Sun, Wen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Expressive Value Learning for Scalable Offline Reinforcement Learning
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
by: Oertell, Owen, et al.
Published: (2024)
by: Oertell, Owen, et al.
Published: (2024)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
by: Liu, Zhongxin, et al.
Published: (2025)
by: Liu, Zhongxin, et al.
Published: (2025)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
by: Bian, Tingcheng, et al.
Published: (2026)
by: Bian, Tingcheng, et al.
Published: (2026)
OER: Offline Experience Replay for Continual Offline Reinforcement Learning
by: Gai, Sibo, et al.
Published: (2023)
by: Gai, Sibo, et al.
Published: (2023)
Offline detection of change-points in the mean for stationary graph signals
by: de la Concha, Alejandro, et al.
Published: (2020)
by: de la Concha, Alejandro, et al.
Published: (2020)
CaMeRL: Collision-Aware and Memory-Enhanced Reinforcement Learning for UAV Navigation in Multi-Scale Obstacle Environments
by: Hong, Hong, et al.
Published: (2026)
by: Hong, Hong, et al.
Published: (2026)
LLMs Can Learn to Reason Via Off-Policy RL
by: Ritter, Daniel, et al.
Published: (2026)
by: Ritter, Daniel, et al.
Published: (2026)
Expressivity of Graph Neural Networks Through the Lens of Adversarial Robustness
by: Campi, Francesco, et al.
Published: (2023)
by: Campi, Francesco, et al.
Published: (2023)
Data-Incremental Continual Offline Reinforcement Learning
by: Gai, Sibo, et al.
Published: (2024)
by: Gai, Sibo, et al.
Published: (2024)
HeadEvolver: Text to Head Avatars via Expressive and Attribute-Preserving Mesh Deformation
by: Wang, Duotun, et al.
Published: (2024)
by: Wang, Duotun, et al.
Published: (2024)
Text2VDM: Text to Vector Displacement Maps for Expressive and Interactive 3D Sculpting
by: Meng, Hengyu, et al.
Published: (2025)
by: Meng, Hengyu, et al.
Published: (2025)
Improving the Expressiveness of $K$-hop Message-Passing GNNs by Injecting Contextualized Substructure Information
by: Yao, Tianjun, et al.
Published: (2024)
by: Yao, Tianjun, et al.
Published: (2024)
Fusing Rewards and Preferences in Reinforcement Learning
by: Khorasani, Sadegh, et al.
Published: (2025)
by: Khorasani, Sadegh, et al.
Published: (2025)
Efficient Action-Constrained Reinforcement Learning via Acceptance-Rejection Method and Augmented MDPs
by: Hung, Wei, et al.
Published: (2025)
by: Hung, Wei, et al.
Published: (2025)
Improving Agent Behaviors with RL Fine-tuning for Autonomous Driving
by: Peng, Zhenghao, et al.
Published: (2024)
by: Peng, Zhenghao, et al.
Published: (2024)
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
by: Chen, Xinjie, et al.
Published: (2026)
by: Chen, Xinjie, et al.
Published: (2026)
OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind
by: Srishty, Sharmin Sultana, et al.
Published: (2026)
by: Srishty, Sharmin Sultana, et al.
Published: (2026)
Diffusing States and Matching Scores: A New Framework for Imitation Learning
by: Wu, Runzhe, et al.
Published: (2024)
by: Wu, Runzhe, et al.
Published: (2024)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
FlowRL: Flow-Augmented Few-Shot Reinforcement Learning for Semi-Structured Sensor Data
by: Pivezhandi, Mohammad, et al.
Published: (2024)
by: Pivezhandi, Mohammad, et al.
Published: (2024)
SIGGesture: Generalized Co-Speech Gesture Synthesis via Semantic Injection with Large-Scale Pre-Training Diffusion Models
by: Cheng, Qingrong, et al.
Published: (2024)
by: Cheng, Qingrong, et al.
Published: (2024)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Semantic State Abstraction Interfaces for LLM-Augmented Portfolio Decisions: Multi-Axis News Decomposition and RL Diagnostics
by: Yerra, Likhita, et al.
Published: (2026)
by: Yerra, Likhita, et al.
Published: (2026)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
by: Tang, Wenjie, et al.
Published: (2026)
by: Tang, Wenjie, et al.
Published: (2026)
Explainable Graph Representation Learning via Graph Pattern Analysis
by: Wang, Xudong, et al.
Published: (2025)
by: Wang, Xudong, et al.
Published: (2025)
ScoRe-Flow: Complete Distributional Control via Score-Based Reinforcement Learning for Flow Matching
by: Qiu, Xiaotian, et al.
Published: (2026)
by: Qiu, Xiaotian, et al.
Published: (2026)
Scaling Laws for Neural Material Models
by: Trikha, Akshay, et al.
Published: (2025)
by: Trikha, Akshay, et al.
Published: (2025)
End-to-end example-based sim-to-real RL policy transfer based on neural stylisation with application to robotic cutting
by: Hathaway, Jamie, et al.
Published: (2026)
by: Hathaway, Jamie, et al.
Published: (2026)
MEReQ: Max-Ent Residual-Q Inverse RL for Sample-Efficient Alignment from Intervention
by: Chen, Yuxin, et al.
Published: (2024)
by: Chen, Yuxin, et al.
Published: (2024)
Dynamic Dual-Granularity Skill Bank for Agentic RL
by: Tu, Songjun, et al.
Published: (2026)
by: Tu, Songjun, et al.
Published: (2026)
The Nash-MTL-STCN Method For Prestack Three-Parameter Inversion
by: Liu, Yingtian, et al.
Published: (2024)
by: Liu, Yingtian, et al.
Published: (2024)
Graceful task adaptation with a bi-hemispheric RL agent
by: Nicholas, Grant, et al.
Published: (2024)
by: Nicholas, Grant, et al.
Published: (2024)
Extending NGU to Multi-Agent RL: A Preliminary Study
by: Hernandez, Juan, et al.
Published: (2025)
by: Hernandez, Juan, et al.
Published: (2025)
On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL
by: Belcamino, Valerio, et al.
Published: (2026)
by: Belcamino, Valerio, et al.
Published: (2026)
Semi-Supervised Learning for AVO Inversion with Strong Spatial Feature Constraints
by: Liu, Yingtian, et al.
Published: (2025)
by: Liu, Yingtian, et al.
Published: (2025)
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
by: Hong, Yoosung
Published: (2026)
by: Hong, Yoosung
Published: (2026)
The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models
by: Liu, Ming
Published: (2026)
by: Liu, Ming
Published: (2026)
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
by: Gray, Gavia, et al.
Published: (2024)
by: Gray, Gavia, et al.
Published: (2024)
Efficient Reinforcement Learning for Global Decision Making in the Presence of Local Agents at Scale
by: Anand, Emile, et al.
Published: (2024)
by: Anand, Emile, et al.
Published: (2024)
Similar Items
-
Expressive Value Learning for Scalable Offline Reinforcement Learning
by: Espinosa-Dice, Nicolas, et al.
Published: (2025) -
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
by: Oertell, Owen, et al.
Published: (2024) -
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
by: Liu, Zhongxin, et al.
Published: (2025) -
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
by: Bian, Tingcheng, et al.
Published: (2026) -
OER: Offline Experience Replay for Continual Offline Reinforcement Learning
by: Gai, Sibo, et al.
Published: (2023)