Gespeichert in:
| Hauptverfasser: | Zhu, Youheng, Lu, Yiping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2603.03191 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recurrent Natural Policy Gradient for POMDPs
von: Cayci, Semih, et al.
Veröffentlicht: (2024)
von: Cayci, Semih, et al.
Veröffentlicht: (2024)
Residuals-based Offline Reinforcement Learning
von: Zhu, Qing, et al.
Veröffentlicht: (2026)
von: Zhu, Qing, et al.
Veröffentlicht: (2026)
Solving Truly Massive Budgeted Monotonic POMDPs with Oracle-Guided Meta-Reinforcement Learning
von: Vora, Manav, et al.
Veröffentlicht: (2024)
von: Vora, Manav, et al.
Veröffentlicht: (2024)
Offline Reinforcement Learning via Inverse Optimization
von: Dimanidis, Ioannis, et al.
Veröffentlicht: (2025)
von: Dimanidis, Ioannis, et al.
Veröffentlicht: (2025)
Dual Control of Linear Systems from Bilinear Observations with Belief Space Model Predictive Control
von: Cao, Daniel, et al.
Veröffentlicht: (2026)
von: Cao, Daniel, et al.
Veröffentlicht: (2026)
Offline Hierarchical Reinforcement Learning via Inverse Optimization
von: Schmidt, Carolin, et al.
Veröffentlicht: (2024)
von: Schmidt, Carolin, et al.
Veröffentlicht: (2024)
Operator Models for Continuous-Time Offline Reinforcement Learning
von: Hoischen, Nicolas, et al.
Veröffentlicht: (2025)
von: Hoischen, Nicolas, et al.
Veröffentlicht: (2025)
On the Width Scaling of Neural Optimizers Under Matrix Operator Norms I: Row/Column Normalization and Hyperparameter Transfer
von: Xu, Ruihan, et al.
Veröffentlicht: (2026)
von: Xu, Ruihan, et al.
Veröffentlicht: (2026)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
von: Zhang, Zhongjun, et al.
Veröffentlicht: (2026)
von: Zhang, Zhongjun, et al.
Veröffentlicht: (2026)
Reward-Relevance-Filtered Linear Offline Reinforcement Learning
von: Zhou, Angela
Veröffentlicht: (2024)
von: Zhou, Angela
Veröffentlicht: (2024)
Online Residual Learning from Offline Experts for Pedestrian Tracking
von: Vlachos, Anastasios, et al.
Veröffentlicht: (2024)
von: Vlachos, Anastasios, et al.
Veröffentlicht: (2024)
Offline Policy Learning with Weight Clipping and Heaviside Composite Optimization
von: Liu, Jingren, et al.
Veröffentlicht: (2026)
von: Liu, Jingren, et al.
Veröffentlicht: (2026)
Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
von: Di, Qiwei, et al.
Veröffentlicht: (2023)
von: Di, Qiwei, et al.
Veröffentlicht: (2023)
PAC-Bayes Meets Online Contextual Optimization
von: Xie, Zhuojun, et al.
Veröffentlicht: (2025)
von: Xie, Zhuojun, et al.
Veröffentlicht: (2025)
Offline Reinforcement Learning via Linear-Programming with Error-Bound Induced Constraints
von: Ozdaglar, Asuman, et al.
Veröffentlicht: (2022)
von: Ozdaglar, Asuman, et al.
Veröffentlicht: (2022)
No-Rank Tensor Decomposition Using Metric Learning
von: Bagherian, Maryam
Veröffentlicht: (2025)
von: Bagherian, Maryam
Veröffentlicht: (2025)
SPP-SBL: Space-Power Prior Sparse Bayesian Learning for Block Sparse Recovery
von: Zhang, Yanhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yanhao, et al.
Veröffentlicht: (2025)
Learning to Cover: Online Learning and Optimization with Irreversible Decisions
von: Jacquillat, Alexandre, et al.
Veröffentlicht: (2024)
von: Jacquillat, Alexandre, et al.
Veröffentlicht: (2024)
SPARKLE: A Unified Single-Loop Primal-Dual Framework for Decentralized Bilevel Optimization
von: Zhu, Shuchen, et al.
Veröffentlicht: (2024)
von: Zhu, Shuchen, et al.
Veröffentlicht: (2024)
OptScaler: A Collaborative Framework for Robust Autoscaling in the Cloud
von: Zou, Ding, et al.
Veröffentlicht: (2023)
von: Zou, Ding, et al.
Veröffentlicht: (2023)
Towards Optimal Offline Reinforcement Learning
von: Li, Mengmeng, et al.
Veröffentlicht: (2025)
von: Li, Mengmeng, et al.
Veröffentlicht: (2025)
Wait-Less Offline Tuning and Re-solving for Online Decision Making
von: Sun, Jingruo, et al.
Veröffentlicht: (2024)
von: Sun, Jingruo, et al.
Veröffentlicht: (2024)
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
von: Zhang, Dake, et al.
Veröffentlicht: (2024)
von: Zhang, Dake, et al.
Veröffentlicht: (2024)
Modeling Hierarchical Spaces: A Review and Unified Framework for Surrogate-Based Architecture Design
von: Saves, Paul, et al.
Veröffentlicht: (2025)
von: Saves, Paul, et al.
Veröffentlicht: (2025)
Unsupervised Ground Metric Learning
von: Auffenberg, Janis, et al.
Veröffentlicht: (2025)
von: Auffenberg, Janis, et al.
Veröffentlicht: (2025)
A Distance Metric for Mixed Integer Programming Instances
von: Maudet, Gwen, et al.
Veröffentlicht: (2025)
von: Maudet, Gwen, et al.
Veröffentlicht: (2025)
Formation Shape Control using the Gromov-Wasserstein Metric
von: Nakashima, Haruto, et al.
Veröffentlicht: (2025)
von: Nakashima, Haruto, et al.
Veröffentlicht: (2025)
Distributionally Robust Optimization via Iterative Algorithms in Continuous Probability Spaces
von: Zhu, Linglingzhi, et al.
Veröffentlicht: (2024)
von: Zhu, Linglingzhi, et al.
Veröffentlicht: (2024)
A Framework for Adaptive Stabilisation of Nonlinear Stochastic Systems
von: Siriya, Seth, et al.
Veröffentlicht: (2025)
von: Siriya, Seth, et al.
Veröffentlicht: (2025)
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
von: Lee, Kyungbok, et al.
Veröffentlicht: (2026)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2026)
Integrated Offline and Online Learning to Solve a Large Class of Scheduling Problems
von: Liu, Anbang, et al.
Veröffentlicht: (2025)
von: Liu, Anbang, et al.
Veröffentlicht: (2025)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
von: Zhang, Weitong, et al.
Veröffentlicht: (2021)
von: Zhang, Weitong, et al.
Veröffentlicht: (2021)
TaskMet: Task-Driven Metric Learning for Model Learning
von: Bansal, Dishank, et al.
Veröffentlicht: (2023)
von: Bansal, Dishank, et al.
Veröffentlicht: (2023)
On the Power of (Approximate) Reward Models for Inference-Time Scaling
von: Zhu, Youheng, et al.
Veröffentlicht: (2026)
von: Zhu, Youheng, et al.
Veröffentlicht: (2026)
A Control Theoretic Framework for Adaptive Gradient Optimizers in Machine Learning
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2022)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2022)
Bisimulation Metrics are Optimal Transport Distances, and Can be Computed Efficiently
von: Calo, Sergio, et al.
Veröffentlicht: (2024)
von: Calo, Sergio, et al.
Veröffentlicht: (2024)
Belief Samples Are All You Need For Social Learning
von: JafariNodeh, Mahyar, et al.
Veröffentlicht: (2024)
von: JafariNodeh, Mahyar, et al.
Veröffentlicht: (2024)
Quantizer Design for Finite Model Approximations, Model Learning, and Quantized Q-Learning for MDPs with Unbounded Spaces
von: Bicer, Osman, et al.
Veröffentlicht: (2025)
von: Bicer, Osman, et al.
Veröffentlicht: (2025)
A Randomized Zeroth-Order Hierarchical Framework for Heterogeneous Federated Learning
von: Qiu, Yuyang, et al.
Veröffentlicht: (2025)
von: Qiu, Yuyang, et al.
Veröffentlicht: (2025)
FERERO: A Flexible Framework for Preference-Guided Multi-Objective Learning
von: Chen, Lisha, et al.
Veröffentlicht: (2024)
von: Chen, Lisha, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Recurrent Natural Policy Gradient for POMDPs
von: Cayci, Semih, et al.
Veröffentlicht: (2024) -
Residuals-based Offline Reinforcement Learning
von: Zhu, Qing, et al.
Veröffentlicht: (2026) -
Solving Truly Massive Budgeted Monotonic POMDPs with Oracle-Guided Meta-Reinforcement Learning
von: Vora, Manav, et al.
Veröffentlicht: (2024) -
Offline Reinforcement Learning via Inverse Optimization
von: Dimanidis, Ioannis, et al.
Veröffentlicht: (2025) -
Dual Control of Linear Systems from Bilinear Observations with Belief Space Model Predictive Control
von: Cao, Daniel, et al.
Veröffentlicht: (2026)