BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement Finetuning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shen, Qianli, Chen, Daoyuan, Huang, Yilun, Ling, Zhenqing, Li, Yaliang, Ding, Bolin, Zhou, Jingren |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data
par: Ling, Zhenqing, et autres
Publié: (2025)
par: Ling, Zhenqing, et autres
Publié: (2025)
Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development
par: Chen, Daoyuan, et autres
Publié: (2024)
par: Chen, Daoyuan, et autres
Publié: (2024)
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
par: Jiao, Qirui, et autres
Publié: (2024)
par: Jiao, Qirui, et autres
Publié: (2024)
Designing Algorithms Empowered by Language Models: An Analytical Framework, Case Studies, and Insights
par: Chen, Yanxi, et autres
Publié: (2024)
par: Chen, Yanxi, et autres
Publié: (2024)
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
par: Xu, Zhe, et autres
Publié: (2025)
par: Xu, Zhe, et autres
Publié: (2025)
Grounded in Reality: Learning and Deploying Proactive LLM from Offline Logs
par: Wei, Fei, et autres
Publié: (2025)
par: Wei, Fei, et autres
Publié: (2025)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
par: Chen, Yanxi, et autres
Publié: (2023)
par: Chen, Yanxi, et autres
Publié: (2023)
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
par: Zhou, Ting, et autres
Publié: (2024)
par: Zhou, Ting, et autres
Publié: (2024)
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
par: Jiao, Qirui, et autres
Publié: (2024)
par: Jiao, Qirui, et autres
Publié: (2024)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
par: Chen, Yanxi, et autres
Publié: (2024)
par: Chen, Yanxi, et autres
Publié: (2024)
EE-Tuning: An Economical yet Scalable Solution for Tuning Early-Exit Large Language Models
par: Pan, Xuchen, et autres
Publié: (2024)
par: Pan, Xuchen, et autres
Publié: (2024)
Tree-based Models for Vertical Federated Learning: A Survey
par: Qian, Bingchen, et autres
Publié: (2025)
par: Qian, Bingchen, et autres
Publié: (2025)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
par: Ge, Ce, et autres
Publié: (2024)
par: Ge, Ce, et autres
Publié: (2024)
The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective
par: Qin, Zhen, et autres
Publié: (2024)
par: Qin, Zhen, et autres
Publié: (2024)
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
par: Jiao, Qirui, et autres
Publié: (2025)
par: Jiao, Qirui, et autres
Publié: (2025)
Towards Anthropomorphic Conversational AI Part I: A Practical Framework
par: Wei, Fei, et autres
Publié: (2025)
par: Wei, Fei, et autres
Publié: (2025)
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
par: Zhang, Wenhao, et autres
Publié: (2025)
par: Zhang, Wenhao, et autres
Publié: (2025)
Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
par: Chen, Daoyuan, et autres
Publié: (2024)
par: Chen, Daoyuan, et autres
Publié: (2024)
Federated Fine-tuning of Large Language Models under Heterogeneous Tasks and Client Resources
par: Bai, Jiamu, et autres
Publié: (2024)
par: Bai, Jiamu, et autres
Publié: (2024)
UniDM: A Unified Framework for Data Manipulation with Large Language Models
par: Qian, Yichen, et autres
Publié: (2024)
par: Qian, Yichen, et autres
Publié: (2024)
On the Convergence of Zeroth-Order Federated Tuning for Large Language Models
par: Ling, Zhenqing, et autres
Publié: (2024)
par: Ling, Zhenqing, et autres
Publié: (2024)
Very Large-Scale Multi-Agent Simulation in AgentScope
par: Pan, Xuchen, et autres
Publié: (2024)
par: Pan, Xuchen, et autres
Publié: (2024)
Output Scaling: YingLong-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model
par: Wang, Xue, et autres
Publié: (2025)
par: Wang, Xue, et autres
Publié: (2025)
Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
par: Ma, Qianli, et autres
Publié: (2024)
par: Ma, Qianli, et autres
Publié: (2024)
Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models
par: Pan, Xuchen, et autres
Publié: (2025)
par: Pan, Xuchen, et autres
Publié: (2025)
BOTS: Batch Bayesian Optimization of Extended Thompson Sampling for Severely Episode-Limited RL Settings
par: Karine, Karine, et autres
Publié: (2024)
par: Karine, Karine, et autres
Publié: (2024)
API-guided Dataset Synthesis to Finetune Large Code Models
par: Li, Zongjie, et autres
Publié: (2024)
par: Li, Zongjie, et autres
Publié: (2024)
An Auction-based Marketplace for Model Trading in Federated Learning
par: Cui, Yue, et autres
Publié: (2024)
par: Cui, Yue, et autres
Publié: (2024)
Towards Fast Safe Online Reinforcement Learning via Policy Finetuning
par: Chen, Keru, et autres
Publié: (2024)
par: Chen, Keru, et autres
Publié: (2024)
Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
par: Zhang, Ling, et autres
Publié: (2025)
par: Zhang, Ling, et autres
Publié: (2025)
A Bargaining-based Approach for Feature Trading in Vertical Federated Learning
par: Cui, Yue, et autres
Publié: (2024)
par: Cui, Yue, et autres
Publié: (2024)
Dynamic Demonstration Retrieval and Cognitive Understanding for Emotional Support Conversation
par: Xu, Zhe, et autres
Publié: (2024)
par: Xu, Zhe, et autres
Publié: (2024)
The Stronger the Diffusion Model, the Easier the Backdoor: Data Poisoning to Induce Copyright Breaches Without Adjusting Finetuning Pipeline
par: Wang, Haonan, et autres
Publié: (2024)
par: Wang, Haonan, et autres
Publié: (2024)
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
par: Duan, Zhongjie, et autres
Publié: (2024)
par: Duan, Zhongjie, et autres
Publié: (2024)
SelectiveFinetuning: Enhancing Transfer Learning in Sleep Staging through Selective Domain Alignment
par: Zhao, Siyuan, et autres
Publié: (2025)
par: Zhao, Siyuan, et autres
Publié: (2025)
On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm
par: Zhou, Zhanpeng, et autres
Publié: (2024)
par: Zhou, Zhanpeng, et autres
Publié: (2024)
Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision Transformers
par: Yan, Kai, et autres
Publié: (2024)
par: Yan, Kai, et autres
Publié: (2024)
Understanding Byzantine Robustness in Federated Learning with A Black-box Server
par: Zhao, Fangyuan, et autres
Publié: (2024)
par: Zhao, Fangyuan, et autres
Publié: (2024)
A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement Learning
par: Guo, Siyuan, et autres
Publié: (2023)
par: Guo, Siyuan, et autres
Publié: (2023)
AgentScope: A Flexible yet Robust Multi-Agent Platform
par: Gao, Dawei, et autres
Publié: (2024)
par: Gao, Dawei, et autres
Publié: (2024)
Documents similaires
-
Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data
par: Ling, Zhenqing, et autres
Publié: (2025) -
Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development
par: Chen, Daoyuan, et autres
Publié: (2024) -
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
par: Jiao, Qirui, et autres
Publié: (2024) -
Designing Algorithms Empowered by Language Models: An Analytical Framework, Case Studies, and Insights
par: Chen, Yanxi, et autres
Publié: (2024) -
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
par: Xu, Zhe, et autres
Publié: (2025)