Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Misra, Dipendra, Pacchiano, Aldo, Chi, Ta-Chung, Gao, Ge
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911401751085056
author Misra, Dipendra
Pacchiano, Aldo
Chi, Ta-Chung
Gao, Ge
author_facet Misra, Dipendra
Pacchiano, Aldo
Chi, Ta-Chung
Gao, Ge
contents We study how to fine-tune LLMs using user-edit deployment data consisting of a set of context, an agent's response, and user edits. This deployment data is naturally generated by users in applications such as LLMs-based writing assistants and coding agents. The _natural_ origin of user edits makes it a desired source for adapting and personalizing LLMs. In this setup, there emerges a unification of various feedback types namely preferences, supervised labels, and cost that are typically studied separately in the literature. In this paper, we initiate the theoretical investigation of learning from user edits. We first derive bounds for learning algorithms that learn from each of these feedback types. We prove that these algorithms have different trade-offs depending upon the user, data distribution, and model class. We then propose a simple ensembling procedure to jointly learn from these feedback types. On two domains adapted from Gao et al. 2024, we show our ensembling procedure outperforms these methods that learn from individual feedback. Further, we show that our proposed procedure can robustly adapt to different user-edit distributions at test time.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19055
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
Misra, Dipendra
Pacchiano, Aldo
Chi, Ta-Chung
Gao, Ge
Machine Learning
Artificial Intelligence
Computation and Language
We study how to fine-tune LLMs using user-edit deployment data consisting of a set of context, an agent's response, and user edits. This deployment data is naturally generated by users in applications such as LLMs-based writing assistants and coding agents. The _natural_ origin of user edits makes it a desired source for adapting and personalizing LLMs. In this setup, there emerges a unification of various feedback types namely preferences, supervised labels, and cost that are typically studied separately in the literature. In this paper, we initiate the theoretical investigation of learning from user edits. We first derive bounds for learning algorithms that learn from each of these feedback types. We prove that these algorithms have different trade-offs depending upon the user, data distribution, and model class. We then propose a simple ensembling procedure to jointly learn from these feedback types. On two domains adapted from Gao et al. 2024, we show our ensembling procedure outperforms these methods that learn from individual feedback. Further, we show that our proposed procedure can robustly adapt to different user-edit distributions at test time.
title Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.19055