Expert or not? assessing data quality in offline reinforcement learning
Fuente:
arXiv
Saved in:
| Main Authors: | Asadulaev, Arip, Karray, Fakhri, Takac, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent Reasoning in TRMs is Secretly a Policy Improvement Operator
by: Asadulaev, Arip, et al.
Published: (2025)
by: Asadulaev, Arip, et al.
Published: (2025)
Y-Shaped Generative Flows
by: Asadulaev, Arip, et al.
Published: (2025)
by: Asadulaev, Arip, et al.
Published: (2025)
Zero-Shot Off-Policy Learning
by: Asadulaev, Arip, et al.
Published: (2026)
by: Asadulaev, Arip, et al.
Published: (2026)
Zero-shot adaptation to order book dynamics
by: Asadulaev, Arip
Published: (2026)
by: Asadulaev, Arip
Published: (2026)
Convex Compositional Reasoning Models
by: Roketlishvili, Meir, et al.
Published: (2026)
by: Roketlishvili, Meir, et al.
Published: (2026)
Light Unbalanced Optimal Transport
by: Gazdieva, Milena, et al.
Published: (2023)
by: Gazdieva, Milena, et al.
Published: (2023)
Neural Optimal Transport with General Cost Functionals
by: Asadulaev, Arip, et al.
Published: (2022)
by: Asadulaev, Arip, et al.
Published: (2022)
Rethinking Optimal Transport in Offline Reinforcement Learning
by: Asadulaev, Arip, et al.
Published: (2024)
by: Asadulaev, Arip, et al.
Published: (2024)
Enhance Hyperbolic Representation Learning via Second-order Pooling
by: Song, Kun, et al.
Published: (2024)
by: Song, Kun, et al.
Published: (2024)
Experimental evaluation of offline reinforcement learning for HVAC control in buildings
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Causal prompting model-based offline reinforcement learning
by: Yu, Xuehui, et al.
Published: (2024)
by: Yu, Xuehui, et al.
Published: (2024)
In-Context Learning Operates as Concept Subspace Learning
by: Tang, Wei, et al.
Published: (2026)
by: Tang, Wei, et al.
Published: (2026)
PyCFRL: A Python library for counterfactually fair offline reinforcement learning via sequential data preprocessing
by: Zhang, Jianhan, et al.
Published: (2025)
by: Zhang, Jianhan, et al.
Published: (2025)
Optimal Transport Maps are Good Voice Converters
by: Asadulaev, Arip, et al.
Published: (2024)
by: Asadulaev, Arip, et al.
Published: (2024)
Deep autoregressive density nets vs neural ensembles for model-based offline reinforcement learning
by: Benechehab, Abdelhakim, et al.
Published: (2024)
by: Benechehab, Abdelhakim, et al.
Published: (2024)
Inverse Entropic Optimal Transport Solves Semi-supervised Learning via Data Likelihood Maximization
by: Persiianov, Mikhail, et al.
Published: (2024)
by: Persiianov, Mikhail, et al.
Published: (2024)
Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
by: Abid, Abderrazek, et al.
Published: (2025)
by: Abid, Abderrazek, et al.
Published: (2025)
Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model
by: Zhang, Jing, et al.
Published: (2024)
by: Zhang, Jing, et al.
Published: (2024)
Physics-informed offline reinforcement learning eliminates catastrophic fuel waste in maritime routing
by: Bora, Aniruddha, et al.
Published: (2026)
by: Bora, Aniruddha, et al.
Published: (2026)
Data-driven simulator of multi-animal behavior with unknown dynamics via offline and online reinforcement learning
by: Fujii, Keisuke, et al.
Published: (2025)
by: Fujii, Keisuke, et al.
Published: (2025)
REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring
by: Ho, Thanh Cong, et al.
Published: (2025)
by: Ho, Thanh Cong, et al.
Published: (2025)
Balancing optimism and pessimism in offline-to-online learning
by: Sentenac, Flore, et al.
Published: (2025)
by: Sentenac, Flore, et al.
Published: (2025)
Bridging Brain with Foundation Models through Self-Supervised Learning
by: Altaheri, Hamdi, et al.
Published: (2025)
by: Altaheri, Hamdi, et al.
Published: (2025)
Adaptive prediction theory combining offline and online learning
by: Li, Haizheng, et al.
Published: (2025)
by: Li, Haizheng, et al.
Published: (2025)
Random-reshuffled SARAH does not need a full gradient computations
by: Beznosikov, Aleksandr, et al.
Published: (2021)
by: Beznosikov, Aleksandr, et al.
Published: (2021)
FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition
by: Islam, Md. Milon, et al.
Published: (2025)
by: Islam, Md. Milon, et al.
Published: (2025)
GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning
by: Raju, S M Taslim Uddin, et al.
Published: (2025)
by: Raju, S M Taslim Uddin, et al.
Published: (2025)
A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition
by: Haque, Md Rezwanul, et al.
Published: (2025)
by: Haque, Md Rezwanul, et al.
Published: (2025)
Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update
by: Li, Qing, et al.
Published: (2025)
by: Li, Qing, et al.
Published: (2025)
Collaborative and Efficient Personalization with Mixtures of Adaptors
by: Almansoori, Abdulla Jasem, et al.
Published: (2024)
by: Almansoori, Abdulla Jasem, et al.
Published: (2024)
Ergodicity in reinforcement learning
by: Baumann, Dominik, et al.
Published: (2026)
by: Baumann, Dominik, et al.
Published: (2026)
Tailoring deep learning for real-time brain-computer interfaces: From offline models to calibration-free online decoding
by: Wimpff, Martin, et al.
Published: (2025)
by: Wimpff, Martin, et al.
Published: (2025)
Regret minimization in Linear Bandits with offline data via extended D-optimal exploration
by: Vijayan, Sushant, et al.
Published: (2025)
by: Vijayan, Sushant, et al.
Published: (2025)
Enhancing BERT Fine-Tuning for Sentiment Analysis in Lower-Resourced Languages
by: Kubík, Jozef, et al.
Published: (2025)
by: Kubík, Jozef, et al.
Published: (2025)
HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
by: Li, Qing, et al.
Published: (2025)
by: Li, Qing, et al.
Published: (2025)
A critical assessment of reinforcement learning methods for microswimmer navigation in complex flows
by: Mecanna, Selim, et al.
Published: (2025)
by: Mecanna, Selim, et al.
Published: (2025)
Generalising Battery Control in Net-Zero Buildings via Personalised Federated RL
by: Avila, Nicolas M Cuadrado, et al.
Published: (2024)
by: Avila, Nicolas M Cuadrado, et al.
Published: (2024)
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
by: Zmushko, Philip, et al.
Published: (2024)
by: Zmushko, Philip, et al.
Published: (2024)
FRESCO: Federated Reinforcement Energy System for Cooperative Optimization
by: Cuadrado, Nicolas Mauricio, et al.
Published: (2024)
by: Cuadrado, Nicolas Mauricio, et al.
Published: (2024)
Byzantine-Robust Optimization under $(L_0, L_1)$-Smoothness
by: Bolatov, Arman, et al.
Published: (2026)
by: Bolatov, Arman, et al.
Published: (2026)
Similar Items
-
Latent Reasoning in TRMs is Secretly a Policy Improvement Operator
by: Asadulaev, Arip, et al.
Published: (2025) -
Y-Shaped Generative Flows
by: Asadulaev, Arip, et al.
Published: (2025) -
Zero-Shot Off-Policy Learning
by: Asadulaev, Arip, et al.
Published: (2026) -
Zero-shot adaptation to order book dynamics
by: Asadulaev, Arip
Published: (2026) -
Convex Compositional Reasoning Models
by: Roketlishvili, Meir, et al.
Published: (2026)