Mechanism Design for LLM Fine-tuning with Multiple Reward Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Haoran, Chen, Yurong, Wang, Siwei, Chu, Xu, Chen, Wei, Deng, Xiaotie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915787169595392
author Sun, Haoran
Chen, Yurong
Wang, Siwei
Chu, Xu
Chen, Wei
Deng, Xiaotie
author_facet Sun, Haoran
Chen, Yurong
Wang, Siwei
Chu, Xu
Chen, Wei
Deng, Xiaotie
contents Fine-tuning large language models (LLMs) to aggregate multiple preferences has attracted considerable research attention. With aggregation algorithms advancing, a potential economic scenario arises where fine-tuning services are provided to agents with different preferences. In this context, agents may benefit from strategically misreporting their preferences, but this could harm the aggregation performance. This paper addresses such incentive issues by framing it as a mechanism design problem: an LLM provider determines the fine-tuning objective (training rule) and the pricing scheme (payment rule) for agents. We primarily focus on training rules that maximize social welfare subject to certain regularizations, referred to as SW-Max rules. First, we show that under most circumstances, truthful reporting is sub-optimal with simply a SW-Max rule, thereby highlighting the necessity of payments. Second, we extend the VCG payment to implement SW-Max rules in dominant-strategy incentive compatibility (DSIC). We characterize sufficient conditions for payment equivalence and derive the necessary conditions for a payment rule to implement a SW-Max rule in DSIC and other principles. Third, we demonstrate that our mechanism is approximately DSIC with perturbed input, showcasing its robustness against the inevitable errors in real-world applications. Experiments on real LLM training results further confirm the practical implications of our results.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16276
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mechanism Design for LLM Fine-tuning with Multiple Reward Models
Sun, Haoran
Chen, Yurong
Wang, Siwei
Chu, Xu
Chen, Wei
Deng, Xiaotie
Computer Science and Game Theory
Fine-tuning large language models (LLMs) to aggregate multiple preferences has attracted considerable research attention. With aggregation algorithms advancing, a potential economic scenario arises where fine-tuning services are provided to agents with different preferences. In this context, agents may benefit from strategically misreporting their preferences, but this could harm the aggregation performance. This paper addresses such incentive issues by framing it as a mechanism design problem: an LLM provider determines the fine-tuning objective (training rule) and the pricing scheme (payment rule) for agents. We primarily focus on training rules that maximize social welfare subject to certain regularizations, referred to as SW-Max rules. First, we show that under most circumstances, truthful reporting is sub-optimal with simply a SW-Max rule, thereby highlighting the necessity of payments. Second, we extend the VCG payment to implement SW-Max rules in dominant-strategy incentive compatibility (DSIC). We characterize sufficient conditions for payment equivalence and derive the necessary conditions for a payment rule to implement a SW-Max rule in DSIC and other principles. Third, we demonstrate that our mechanism is approximately DSIC with perturbed input, showcasing its robustness against the inevitable errors in real-world applications. Experiments on real LLM training results further confirm the practical implications of our results.
title Mechanism Design for LLM Fine-tuning with Multiple Reward Models
topic Computer Science and Game Theory
url https://arxiv.org/abs/2405.16276