Saved in:
Bibliographic Details
Main Authors: Ferragu, Constance, Ziegler, Jonathan D., Deutschmann, Nicolas, Lindoulsi, Arthur, Bixby, Eli, Team, Cradle ML
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.19474
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911288039309312
author Ferragu, Constance
Ziegler, Jonathan D.
Deutschmann, Nicolas
Lindoulsi, Arthur
Bixby, Eli
Team, Cradle ML
author_facet Ferragu, Constance
Ziegler, Jonathan D.
Deutschmann, Nicolas
Lindoulsi, Arthur
Bixby, Eli
Team, Cradle ML
contents Direct Preference Optimization (DPO) is an effective approach for aligning protein language models with experimental design goals. However, DPO faces a scalability bottleneck: the number of possible training pairs grows quadratically with the number of labeled sequences, leading to prohibitive training times even for modestly sized datasets. We introduce g-DPO, a framework that (i) uses sequence space clustering to prune redundant pairs while preserving training signal, and (ii) amortizes likelihood computations with group-based approximations. Across three protein engineering tasks, g-DPO maintains in silico and in vitro performance that is statistically indistinguishable from standard DPO, while converging 1.7x to 5.4x times faster, with speedups that scale with dataset size and the structure of the underlying mutational landscape.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19474
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle g-DPO: Scalable Preference Optimization for Protein Language Models
Ferragu, Constance
Ziegler, Jonathan D.
Deutschmann, Nicolas
Lindoulsi, Arthur
Bixby, Eli
Team, Cradle ML
Machine Learning
Direct Preference Optimization (DPO) is an effective approach for aligning protein language models with experimental design goals. However, DPO faces a scalability bottleneck: the number of possible training pairs grows quadratically with the number of labeled sequences, leading to prohibitive training times even for modestly sized datasets. We introduce g-DPO, a framework that (i) uses sequence space clustering to prune redundant pairs while preserving training signal, and (ii) amortizes likelihood computations with group-based approximations. Across three protein engineering tasks, g-DPO maintains in silico and in vitro performance that is statistically indistinguishable from standard DPO, while converging 1.7x to 5.4x times faster, with speedups that scale with dataset size and the structure of the underlying mutational landscape.
title g-DPO: Scalable Preference Optimization for Protein Language Models
topic Machine Learning
url https://arxiv.org/abs/2510.19474