Demystifying Design Choices of Reinforcement Fine-tuning: A Batched Contextual Bandit Learning Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Hong, Hu, Xiao, Tan, Tao, Gu, Haoran, Li, Xin, Han, Jianyu, Lian, Defu, Chen, Enhong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!

Similar Items