Multiple Choice Learning of Low-Rank Adapters for Language Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Letzelter, Victor, Malard, Hugo, Fontaine, Mathieu, Richard, Gaël, Essid, Slim, Bursuc, Andrei, Pérez, Patrick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918321744510976
author Letzelter, Victor
Malard, Hugo
Fontaine, Mathieu
Richard, Gaël
Essid, Slim
Bursuc, Andrei
Pérez, Patrick
author_facet Letzelter, Victor
Malard, Hugo
Fontaine, Mathieu
Richard, Gaël
Essid, Slim
Bursuc, Andrei
Pérez, Patrick
contents We propose LoRA-MCL, a training scheme that extends next-token prediction in language models with a method designed to decode diverse, plausible sentence continuations at inference time. Traditional language modeling is an intrinsically ill-posed problem: given a context, multiple ``futures'' may be equally plausible. Our approach leverages Multiple Choice Learning (MCL) and the Winner-Takes-All loss to efficiently handle ambiguity through Low-Rank Adaptation. We provide a theoretical interpretation of applying MCL to language modeling, assuming the data is generated from a mixture of distributions. We illustrate the proposed approach using mixtures of Markov chains. We then demonstrate with experiments on visual and audio captioning, as well as machine translation, that our method achieves high diversity and relevance in generated outputs. The accompanying code and a general-purpose package for applying LoRA-MCL to a wide range of language models are made available.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10419
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multiple Choice Learning of Low-Rank Adapters for Language Modeling
Letzelter, Victor
Malard, Hugo
Fontaine, Mathieu
Richard, Gaël
Essid, Slim
Bursuc, Andrei
Pérez, Patrick
Machine Learning
Artificial Intelligence
Computation and Language
We propose LoRA-MCL, a training scheme that extends next-token prediction in language models with a method designed to decode diverse, plausible sentence continuations at inference time. Traditional language modeling is an intrinsically ill-posed problem: given a context, multiple ``futures'' may be equally plausible. Our approach leverages Multiple Choice Learning (MCL) and the Winner-Takes-All loss to efficiently handle ambiguity through Low-Rank Adaptation. We provide a theoretical interpretation of applying MCL to language modeling, assuming the data is generated from a mixture of distributions. We illustrate the proposed approach using mixtures of Markov chains. We then demonstrate with experiments on visual and audio captioning, as well as machine translation, that our method achieves high diversity and relevance in generated outputs. The accompanying code and a general-purpose package for applying LoRA-MCL to a wide range of language models are made available.
title Multiple Choice Learning of Low-Rank Adapters for Language Modeling
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.10419