Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghosh, Satanu, Holgate, Collin, Brodnik, Neal R., Downey, Doug, Daly, Samantha, Pollock, Tresa M., Carton, Samuel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918202501496832
author Ghosh, Satanu
Holgate, Collin
Brodnik, Neal R.
Downey, Doug
Daly, Samantha
Pollock, Tresa M.
Carton, Samuel
author_facet Ghosh, Satanu
Holgate, Collin
Brodnik, Neal R.
Downey, Doug
Daly, Samantha
Pollock, Tresa M.
Carton, Samuel
contents We apply preference learning to the task of language model-guided design of novel structural alloys. In contrast to prior work that focuses on generating stable inorganic crystals, our approach targets the synthesizeability of a specific structural class: BCC/B2 superalloys, an underexplored family of materials with potential applications in extreme environments. Using three open-weight models (LLaMA-3.1, Gemma-2, and OLMo-2), we demonstrate that language models can be optimized for multiple design objectives using a single, unified reward signal through Direct Preference Optimization (DPO). Unlike prior approaches that rely on heuristic or human-in-the-loop feedback (costly), our reward signal is derived from thermodynamic phase calculations, offering a scientifically grounded criterion for model tuning. To our knowledge, this is the first demonstration of preference-tuning a language model using physics-grounded feedback for structural alloy design. The resulting framework is general and extensible, providing a path forward for intelligent design-space exploration across a range of physical science domains.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12036
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys
Ghosh, Satanu
Holgate, Collin
Brodnik, Neal R.
Downey, Doug
Daly, Samantha
Pollock, Tresa M.
Carton, Samuel
Computational Engineering, Finance, and Science
Materials Science
Artificial Intelligence
Computation and Language
Machine Learning
80, 92
I.2.m; I.2.7; J.2
We apply preference learning to the task of language model-guided design of novel structural alloys. In contrast to prior work that focuses on generating stable inorganic crystals, our approach targets the synthesizeability of a specific structural class: BCC/B2 superalloys, an underexplored family of materials with potential applications in extreme environments. Using three open-weight models (LLaMA-3.1, Gemma-2, and OLMo-2), we demonstrate that language models can be optimized for multiple design objectives using a single, unified reward signal through Direct Preference Optimization (DPO). Unlike prior approaches that rely on heuristic or human-in-the-loop feedback (costly), our reward signal is derived from thermodynamic phase calculations, offering a scientifically grounded criterion for model tuning. To our knowledge, this is the first demonstration of preference-tuning a language model using physics-grounded feedback for structural alloy design. The resulting framework is general and extensible, providing a path forward for intelligent design-space exploration across a range of physical science domains.
title Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys
topic Computational Engineering, Finance, and Science
Materials Science
Artificial Intelligence
Computation and Language
Machine Learning
80, 92
I.2.m; I.2.7; J.2
url https://arxiv.org/abs/2511.12036