On the Role of Iterative Computation in Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Ghugare, Raj, Bortkiewicz, Michał, Ziarko, Alicja, Eysenbach, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contrastive Representations for Temporal Reasoning
by: Ziarko, Alicja, et al.
Published: (2025)
by: Ziarko, Alicja, et al.
Published: (2025)
Normalizing Flows are Capable Models for RL
by: Ghugare, Raj, et al.
Published: (2025)
by: Ghugare, Raj, et al.
Published: (2025)
Closing the Gap between TD Learning and Supervised Learning -- A Generalisation Point of View
by: Ghugare, Raj, et al.
Published: (2024)
by: Ghugare, Raj, et al.
Published: (2024)
Is Temporal Difference Learning the Gold Standard for Stitching in RL?
by: Bortkiewicz, Michał, et al.
Published: (2025)
by: Bortkiewicz, Michał, et al.
Published: (2025)
1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
by: Wang, Kevin, et al.
Published: (2025)
by: Wang, Kevin, et al.
Published: (2025)
Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
by: Góral, Gracjan, et al.
Published: (2024)
by: Góral, Gracjan, et al.
Published: (2024)
BuilderBench: The Building Blocks of Intelligent Agents
by: Ghugare, Raj, et al.
Published: (2025)
by: Ghugare, Raj, et al.
Published: (2025)
Accelerating Goal-Conditioned RL Algorithms and Research
by: Bortkiewicz, Michał, et al.
Published: (2024)
by: Bortkiewicz, Michał, et al.
Published: (2024)
Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning
by: Surdej, Rafał, et al.
Published: (2025)
by: Surdej, Rafał, et al.
Published: (2025)
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Repurposing Language Models into Embedding Models: Finding the Compute-Optimal Recipe
by: Ziarko, Alicja, et al.
Published: (2024)
by: Ziarko, Alicja, et al.
Published: (2024)
Horizon Generalization in Reinforcement Learning
by: Myers, Vivek, et al.
Published: (2025)
by: Myers, Vivek, et al.
Published: (2025)
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
by: Wołczyk, Maciej, et al.
Published: (2024)
by: Wołczyk, Maciej, et al.
Published: (2024)
Offline Goal-conditioned Reinforcement Learning with Quasimetric Representations
by: Myers, Vivek, et al.
Published: (2025)
by: Myers, Vivek, et al.
Published: (2025)
Multistep Quasimetric Learning for Scalable Goal-conditioned Reinforcement Learning
by: Zheng, Bill Chunyuan, et al.
Published: (2025)
by: Zheng, Bill Chunyuan, et al.
Published: (2025)
Learning to Perceive the World Through Control: Empowerment-Based Representation Learning
by: Bastankhah, Mahsa, et al.
Published: (2026)
by: Bastankhah, Mahsa, et al.
Published: (2026)
The "Law" of the Unconscious Contrastive Learner: Probabilistic Alignment of Unpaired Modalities
by: Che, Yongwei, et al.
Published: (2025)
by: Che, Yongwei, et al.
Published: (2025)
Behavior-Consistent Deep Reinforcement Learning
by: Hussing, Marcel, et al.
Published: (2026)
by: Hussing, Marcel, et al.
Published: (2026)
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning
by: Venugopal, Aravind, et al.
Published: (2026)
by: Venugopal, Aravind, et al.
Published: (2026)
Learning Continually by Spectral Regularization
by: Lewandowski, Alex, et al.
Published: (2024)
by: Lewandowski, Alex, et al.
Published: (2024)
Consistent Zero-Shot Imitation with Contrastive Goal Inference
by: Wantlin, Kathryn, et al.
Published: (2025)
by: Wantlin, Kathryn, et al.
Published: (2025)
Can We Really Learn One Representation to Optimize All Rewards?
by: Zheng, Chongyi, et al.
Published: (2026)
by: Zheng, Chongyi, et al.
Published: (2026)
SLAP: Shortcut Learning for Abstract Planning
by: Liu, Y. Isabel, et al.
Published: (2025)
by: Liu, Y. Isabel, et al.
Published: (2025)
Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic Rewards
by: Mohamed, Faisal, et al.
Published: (2026)
by: Mohamed, Faisal, et al.
Published: (2026)
Can a MISL Fly? Analysis and Ingredients for Mutual Information Skill Learning
by: Zheng, Chongyi, et al.
Published: (2024)
by: Zheng, Chongyi, et al.
Published: (2024)
A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals
by: Liu, Grace, et al.
Published: (2024)
by: Liu, Grace, et al.
Published: (2024)
Contrastive Difference Predictive Coding
by: Zheng, Chongyi, et al.
Published: (2023)
by: Zheng, Chongyi, et al.
Published: (2023)
Inference via Interpolation: Contrastive Representations Provably Enable Planning and Inference
by: Eysenbach, Benjamin, et al.
Published: (2024)
by: Eysenbach, Benjamin, et al.
Published: (2024)
Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled Perturbations
by: Liang, Yongyuan, et al.
Published: (2023)
by: Liang, Yongyuan, et al.
Published: (2023)
Unifying Goal-Conditioned RL and Unsupervised Skill Learning via Control-Maximization
by: Modirshanechi, Alireza, et al.
Published: (2026)
by: Modirshanechi, Alireza, et al.
Published: (2026)
UpSkill: Mutual Information Skill Learning for Structured Response Diversity in LLMs
by: Shah, Devan, et al.
Published: (2026)
by: Shah, Devan, et al.
Published: (2026)
Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making
by: Myers, Vivek, et al.
Published: (2024)
by: Myers, Vivek, et al.
Published: (2024)
A Rate-Distortion View of Uncertainty Quantification
by: Apostolopoulou, Ifigeneia, et al.
Published: (2024)
by: Apostolopoulou, Ifigeneia, et al.
Published: (2024)
Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration
by: Nimonkar, Chirayu, et al.
Published: (2025)
by: Nimonkar, Chirayu, et al.
Published: (2025)
OGBench: Benchmarking Offline Goal-Conditioned RL
by: Park, Seohong, et al.
Published: (2024)
by: Park, Seohong, et al.
Published: (2024)
Intention-Conditioned Flow Occupancy Models
by: Zheng, Chongyi, et al.
Published: (2025)
by: Zheng, Chongyi, et al.
Published: (2025)
HIQL: Offline Goal-Conditioned RL with Latent States as Actions
by: Park, Seohong, et al.
Published: (2023)
by: Park, Seohong, et al.
Published: (2023)
Linear Feedback Control Systems for Iterative Prompt Optimization in Large Language Models
by: Karn, Rupesh Raj
Published: (2025)
by: Karn, Rupesh Raj
Published: (2025)
Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL
by: Bastankhah, Mahsa, et al.
Published: (2025)
by: Bastankhah, Mahsa, et al.
Published: (2025)
Similar Items
-
Contrastive Representations for Temporal Reasoning
by: Ziarko, Alicja, et al.
Published: (2025) -
Normalizing Flows are Capable Models for RL
by: Ghugare, Raj, et al.
Published: (2025) -
Closing the Gap between TD Learning and Supervised Learning -- A Generalisation Point of View
by: Ghugare, Raj, et al.
Published: (2024) -
Is Temporal Difference Learning the Gold Standard for Stitching in RL?
by: Bortkiewicz, Michał, et al.
Published: (2025) -
1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
by: Wang, Kevin, et al.
Published: (2025)