Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Duong, Thang, Yang, Minglai, Zhang, Chicheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Task Diversity: Provable Representation Transfer for Sequential Multi-Task Linear Bandits
by: Duong, Thang, et al.
Published: (2025)
by: Duong, Thang, et al.
Published: (2025)
Physics-Informed Parametric Bandits for Beam Alignment in mmWave Communications
by: Qin, Hao, et al.
Published: (2025)
by: Qin, Hao, et al.
Published: (2025)
Towards Fundamental Limits for Active Multi-distribution Learning
by: Zhang, Chicheng, et al.
Published: (2025)
by: Zhang, Chicheng, et al.
Published: (2025)
Agnostic Interactive Imitation Learning: New Theory and Practical Algorithms
by: Li, Yichen, et al.
Published: (2023)
by: Li, Yichen, et al.
Published: (2023)
Interactive and Hybrid Imitation Learning: Provably Beating Behavior Cloning
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
Efficient Active Learning Halfspaces with Tsybakov Noise: A Non-convex Optimization Approach
by: Li, Yinan, et al.
Published: (2023)
by: Li, Yinan, et al.
Published: (2023)
Taming the Monster Every Context: Complexity Measure and Unified Framework for Offline-Oracle Efficient Contextual Bandits
by: Qin, Hao, et al.
Published: (2026)
by: Qin, Hao, et al.
Published: (2026)
Bridging Lifelong and Multi-Task Representation Learning via Algorithm and Complexity Measure
by: Wang, Zhi, et al.
Published: (2025)
by: Wang, Zhi, et al.
Published: (2025)
LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment
by: Li, Shipeng, et al.
Published: (2025)
by: Li, Shipeng, et al.
Published: (2025)
Word2VecGD: Neural Graph Drawing with Cosine-Stress Optimization
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits
by: Jang, Kyoungseok, et al.
Published: (2024)
by: Jang, Kyoungseok, et al.
Published: (2024)
Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
by: Qin, Hao, et al.
Published: (2023)
by: Qin, Hao, et al.
Published: (2023)
Improving Fairness in Graph Neural Networks via Counterfactual Debiasing
by: Wo, Zengyi, et al.
Published: (2025)
by: Wo, Zengyi, et al.
Published: (2025)
An Extremely Data-efficient and Generative LLM-based Reinforcement Learning Agent for Recommenders
by: Feng, Shuang, et al.
Published: (2024)
by: Feng, Shuang, et al.
Published: (2024)
Warm-starting Push-Relabel
by: Davies, Sami, et al.
Published: (2024)
by: Davies, Sami, et al.
Published: (2024)
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling
by: Qin, Hao, et al.
Published: (2025)
by: Qin, Hao, et al.
Published: (2025)
Controllable Expensive Multi-objective Learning with Warm-starting Bayesian Optimization
by: Nguyen, Quang-Huy, et al.
Published: (2023)
by: Nguyen, Quang-Huy, et al.
Published: (2023)
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
Statistical Inference for Sequential Feature Selection after Domain Adaptation
by: Loc, Duong Tan, et al.
Published: (2025)
by: Loc, Duong Tan, et al.
Published: (2025)
Rethinking of Encoder-based Warm-start Methods in Hyperparameter Optimization
by: Płudowski, Dawid, et al.
Published: (2024)
by: Płudowski, Dawid, et al.
Published: (2024)
Statistical Inference for Feature Selection after Optimal Transport-based Domain Adaptation
by: Loi, Nguyen Thang, et al.
Published: (2024)
by: Loi, Nguyen Thang, et al.
Published: (2024)
EchoRL: Reinforcement Learning via Rollout Echoing
by: Bi, Jinhe, et al.
Published: (2026)
by: Bi, Jinhe, et al.
Published: (2026)
Evolutionary Warm-Starts for Reinforcement Learning in Industrial Continuous Control
by: Maus, Tom, et al.
Published: (2026)
by: Maus, Tom, et al.
Published: (2026)
Warm-starting active-set solvers using graph neural networks
by: Schmidtobreick, Ella J., et al.
Published: (2025)
by: Schmidtobreick, Ella J., et al.
Published: (2025)
Reinforcement Learning for Causal Discovery without Acyclicity Constraints
by: Duong, Bao, et al.
Published: (2024)
by: Duong, Bao, et al.
Published: (2024)
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
by: Yang, Ningyuan, et al.
Published: (2026)
by: Yang, Ningyuan, et al.
Published: (2026)
Learning Fair Invariant Representations under Covariate and Correlation Shifts Simultaneously
by: Li, Dong, et al.
Published: (2024)
by: Li, Dong, et al.
Published: (2024)
Graph Bayesian Optimization for Multiplex Influence Maximization
by: Yuan, Zirui, et al.
Published: (2024)
by: Yuan, Zirui, et al.
Published: (2024)
TabReason: A Reinforcement Learning-Enhanced Reasoning LLM for Explainable Tabular Data Prediction
by: Xu, Tommy, et al.
Published: (2025)
by: Xu, Tommy, et al.
Published: (2025)
Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
by: Nguyen, Khoi Anh, et al.
Published: (2025)
by: Nguyen, Khoi Anh, et al.
Published: (2025)
Improving Iterative Gaussian Processes via Warm Starting Sequential Posteriors
by: Dong, Alan Yufei, et al.
Published: (2025)
by: Dong, Alan Yufei, et al.
Published: (2025)
Data Acquisition for Improving Model Fairness using Reinforcement Learning
by: Hasan, Jahid, et al.
Published: (2024)
by: Hasan, Jahid, et al.
Published: (2024)
LACONIC: Length-Aware Constrained Reinforcement Learning for LLM
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
AGILE: A Novel Reinforcement Learning Framework of LLM Agents
by: Feng, Peiyuan, et al.
Published: (2024)
by: Feng, Peiyuan, et al.
Published: (2024)
Improving Generative Ad Text on Facebook using Reinforcement Learning
by: Jiang, Daniel R., et al.
Published: (2025)
by: Jiang, Daniel R., et al.
Published: (2025)
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
by: Parashar, Shubham, et al.
Published: (2025)
by: Parashar, Shubham, et al.
Published: (2025)
MLDGG: Meta-Learning for Domain Generalization on Graphs
by: Tian, Qin, et al.
Published: (2024)
by: Tian, Qin, et al.
Published: (2024)
Adaptive Graph Mixture of Residual Experts: Unsupervised Learning on Diverse Graphs with Heterogeneous Specialization
by: Chu, Yunlong, et al.
Published: (2025)
by: Chu, Yunlong, et al.
Published: (2025)
Any-step Dynamics Model Improves Future Predictions for Online and Offline Reinforcement Learning
by: Lin, Haoxin, et al.
Published: (2024)
by: Lin, Haoxin, et al.
Published: (2024)
Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning
by: Yan, John, et al.
Published: (2026)
by: Yan, John, et al.
Published: (2026)
Similar Items
-
Beyond Task Diversity: Provable Representation Transfer for Sequential Multi-Task Linear Bandits
by: Duong, Thang, et al.
Published: (2025) -
Physics-Informed Parametric Bandits for Beam Alignment in mmWave Communications
by: Qin, Hao, et al.
Published: (2025) -
Towards Fundamental Limits for Active Multi-distribution Learning
by: Zhang, Chicheng, et al.
Published: (2025) -
Agnostic Interactive Imitation Learning: New Theory and Practical Algorithms
by: Li, Yichen, et al.
Published: (2023) -
Interactive and Hybrid Imitation Learning: Provably Beating Behavior Cloning
by: Li, Yichen, et al.
Published: (2024)