Federated Fine-Tuning of Large Language Models: Kahneman-Tversky vs. Direct Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Spadea, Fernando, Seneviratne, Oshani
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909502408753152
author Spadea, Fernando
Seneviratne, Oshani
author_facet Spadea, Fernando
Seneviratne, Oshani
contents We evaluate Kahneman-Tversky Optimization (KTO) as a fine-tuning method for large language models (LLMs) in federated learning (FL) settings, comparing it against Direct Preference Optimization (DPO). Using Alpaca-7B as the base model, we fine-tune on a realistic dataset under both methods and evaluate performance using MT-Bench-1, Vicuna, and AdvBench benchmarks. Additionally, we introduce a redistributed dataset setup, where only KTO is applicable due to its ability to handle single-response feedback, unlike DPO's reliance on paired responses. Our results demonstrate that KTO, in both its original (KTOO) and redistributed (KTOR) configurations, consistently outperforms DPO across all benchmarks. In the redistributed setup, KTO further validates its flexibility and resilience by maintaining superior performance in scenarios where DPO cannot be applied. These findings establish KTO as a robust and scalable fine-tuning method for FL, motivating its adoption for privacy-preserving, decentralized, and heterogeneous environments.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14187
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Federated Fine-Tuning of Large Language Models: Kahneman-Tversky vs. Direct Preference Optimization
Spadea, Fernando
Seneviratne, Oshani
Machine Learning
Computation and Language
We evaluate Kahneman-Tversky Optimization (KTO) as a fine-tuning method for large language models (LLMs) in federated learning (FL) settings, comparing it against Direct Preference Optimization (DPO). Using Alpaca-7B as the base model, we fine-tune on a realistic dataset under both methods and evaluate performance using MT-Bench-1, Vicuna, and AdvBench benchmarks. Additionally, we introduce a redistributed dataset setup, where only KTO is applicable due to its ability to handle single-response feedback, unlike DPO's reliance on paired responses. Our results demonstrate that KTO, in both its original (KTOO) and redistributed (KTOR) configurations, consistently outperforms DPO across all benchmarks. In the redistributed setup, KTO further validates its flexibility and resilience by maintaining superior performance in scenarios where DPO cannot be applied. These findings establish KTO as a robust and scalable fine-tuning method for FL, motivating its adoption for privacy-preserving, decentralized, and heterogeneous environments.
title Federated Fine-Tuning of Large Language Models: Kahneman-Tversky vs. Direct Preference Optimization
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2502.14187