RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Arzhantsev, Aleksei, Sakhi, Otmane, Vasile, Flavian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Consistency via Marginal Sharpening
by: Arzhantsev, Aleksei, et al.
Published: (2026)
by: Arzhantsev, Aleksei, et al.
Published: (2026)
RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
by: Aouali, Imad, et al.
Published: (2026)
by: Aouali, Imad, et al.
Published: (2026)
Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation
by: Aouali, Imad, et al.
Published: (2022)
by: Aouali, Imad, et al.
Published: (2022)
Sequential Off-Policy Learning with Logarithmic Smoothing
by: Haddouche, Maxime, et al.
Published: (2025)
by: Haddouche, Maxime, et al.
Published: (2025)
Learning to Bid in Repeated Second-Price Auctions with Dynamic Values and Aggregated Feedback
by: Heymann, Benjamin, et al.
Published: (2026)
by: Heymann, Benjamin, et al.
Published: (2026)
Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
by: Aouali, Imad, et al.
Published: (2025)
by: Aouali, Imad, et al.
Published: (2025)
Offline Contextual Bandit with Counterfactual Sample Identification
by: Gilotte, Alexandre, et al.
Published: (2025)
by: Gilotte, Alexandre, et al.
Published: (2025)
Non-Linear Counterfactual Aggregate Optimization
by: Heymann, Benjamin, et al.
Published: (2025)
by: Heymann, Benjamin, et al.
Published: (2025)
Data Valuation for LLM Fine-Tuning: Efficient Shapley Value Approximation via Language Model Arithmetic
by: Tamine, Mélissa, et al.
Published: (2025)
by: Tamine, Mélissa, et al.
Published: (2025)
Exploiting Similarities in A/B Testing with Off-Policy Estimation
by: Sakhi, Otmane, et al.
Published: (2025)
by: Sakhi, Otmane, et al.
Published: (2025)
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
by: Sakhi, Otmane, et al.
Published: (2024)
by: Sakhi, Otmane, et al.
Published: (2024)
Fast Slate Policy Optimization: Going Beyond Plackett-Luce
by: Sakhi, Otmane, et al.
Published: (2023)
by: Sakhi, Otmane, et al.
Published: (2023)
Survival Reinforcement Learning: Toward Scalable Self-Supervised RL
by: Nguimatsia-Tiofack, Franki, et al.
Published: (2026)
by: Nguimatsia-Tiofack, Franki, et al.
Published: (2026)
RL-BioAug: Label-Efficient Reinforcement Learning for Self-Supervised EEG Representation Learning
by: Lee, Cheol-Hui, et al.
Published: (2026)
by: Lee, Cheol-Hui, et al.
Published: (2026)
Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
by: Wang, Qi, et al.
Published: (2023)
by: Wang, Qi, et al.
Published: (2023)
Algorithmic Guarantees for Distilling Supervised and Offline RL Datasets
by: Gupta, Aaryan, et al.
Published: (2025)
by: Gupta, Aaryan, et al.
Published: (2025)
Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning
by: Liao, Luofeng, et al.
Published: (2021)
by: Liao, Luofeng, et al.
Published: (2021)
Sample-Efficient Tabular Self-Play for Offline Robust Reinforcement Learning
by: Li, Na, et al.
Published: (2025)
by: Li, Na, et al.
Published: (2025)
Adaptive $Q$-Aid for Conditional Supervised Learning in Offline Reinforcement Learning
by: Kim, Jeonghye, et al.
Published: (2024)
by: Kim, Jeonghye, et al.
Published: (2024)
Agentic Planning with Reasoning for Image Styling via Offline RL
by: Mukherjee, Subhojyoti, et al.
Published: (2026)
by: Mukherjee, Subhojyoti, et al.
Published: (2026)
Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation
by: Li, Huanyu, et al.
Published: (2026)
by: Li, Huanyu, et al.
Published: (2026)
REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning
by: Deng, Hexuan, et al.
Published: (2025)
by: Deng, Hexuan, et al.
Published: (2025)
Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Diffusion Self-Weighted Guidance for Offline Reinforcement Learning
by: Tagle, Augusto, et al.
Published: (2025)
by: Tagle, Augusto, et al.
Published: (2025)
Boundary-to-Region Supervision for Offline Safe Reinforcement Learning
by: Su, Huikang, et al.
Published: (2025)
by: Su, Huikang, et al.
Published: (2025)
VendiRL: A Framework for Self-Supervised Reinforcement Learning of Diversely Diverse Skills
by: Lintunen, Erik M.
Published: (2025)
by: Lintunen, Erik M.
Published: (2025)
Efficient Offline Reinforcement Learning: First Imitate, then Improve
by: Jelley, Adam, et al.
Published: (2024)
by: Jelley, Adam, et al.
Published: (2024)
Exploring 3D-aware Latent Spaces for Efficiently Learning Numerous Scenes
by: Schnepf, Antoine, et al.
Published: (2024)
by: Schnepf, Antoine, et al.
Published: (2024)
Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
by: Bai, Chenjia, et al.
Published: (2024)
by: Bai, Chenjia, et al.
Published: (2024)
CtRL-Sim: Reactive and Controllable Driving Agents with Offline Reinforcement Learning
by: Rowe, Luke, et al.
Published: (2024)
by: Rowe, Luke, et al.
Published: (2024)
SelfBC: Self Behavior Cloning for Offline Reinforcement Learning
by: Liu, Shirong, et al.
Published: (2024)
by: Liu, Shirong, et al.
Published: (2024)
RL-STaR: Theoretical Analysis of Reinforcement Learning Frameworks for Self-Taught Reasoner
by: Chang, Fu-Chieh, et al.
Published: (2024)
by: Chang, Fu-Chieh, et al.
Published: (2024)
SCOPE-RL: A Python Library for Offline Reinforcement Learning and Off-Policy Evaluation
by: Kiyohara, Haruka, et al.
Published: (2023)
by: Kiyohara, Haruka, et al.
Published: (2023)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
by: Li, Yuhang, et al.
Published: (2026)
by: Li, Yuhang, et al.
Published: (2026)
Offline Inverse RL: New Solution Concepts and Provably Efficient Algorithms
by: Lazzati, Filippo, et al.
Published: (2024)
by: Lazzati, Filippo, et al.
Published: (2024)
Contrastive UCB: Provably Efficient Contrastive Self-Supervised Learning in Online Reinforcement Learning
by: Qiu, Shuang, et al.
Published: (2022)
by: Qiu, Shuang, et al.
Published: (2022)
Self-Supervised Learning of Iterative Solvers for Constrained Optimization
by: Lüken, Lukas, et al.
Published: (2024)
by: Lüken, Lukas, et al.
Published: (2024)
Diffusion Actor-Critic: Formulating Constrained Policy Iteration as Diffusion Noise Regression for Offline Reinforcement Learning
by: Fang, Linjiajie, et al.
Published: (2024)
by: Fang, Linjiajie, et al.
Published: (2024)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning
by: Chaudhary, Gaurav, et al.
Published: (2025)
by: Chaudhary, Gaurav, et al.
Published: (2025)
Similar Items
-
Self-Consistency via Marginal Sharpening
by: Arzhantsev, Aleksei, et al.
Published: (2026) -
RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
by: Aouali, Imad, et al.
Published: (2026) -
Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation
by: Aouali, Imad, et al.
Published: (2022) -
Sequential Off-Policy Learning with Logarithmic Smoothing
by: Haddouche, Maxime, et al.
Published: (2025) -
Learning to Bid in Repeated Second-Price Auctions with Dynamic Values and Aggregated Feedback
by: Heymann, Benjamin, et al.
Published: (2026)