PLaID++: A Preference Aligned Language Model for Targeted Inorganic Materials Design

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Andy, Desai, Rohan, Wang, Larry, Hope, Gabriel, Ritz, Ethan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911328727203840
author Xu, Andy
Desai, Rohan
Wang, Larry
Hope, Gabriel
Ritz, Ethan
author_facet Xu, Andy
Desai, Rohan
Wang, Larry
Hope, Gabriel
Ritz, Ethan
contents Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising approach to improve correctness in LLMs, however, in many scientific problems, the objective is not necessarily to produce the correct answer, but instead to produce a diverse array of candidates which satisfy a set of constraints. We study this challenge in the context of materials generation. To this end, we introduce PLaID++, an LLM post-trained for stable and property-guided crystal generation. We find that performance hinges on our crystallographic representation and reward formulation. First, we introduce a compact, symmetry-informed Wyckoff text representation which improves computational efficiency and encourages generalization from physical priors. Second, we demonstrate that temperature scaling acts as an entropy regularizer which counteracts mode collapse and encourages exploration. By encoding symmetry constraints directly into text and guiding model outputs towards desirable chemical space, PLaID++ generates structures that are thermodynamically stable, unique, and novel at a $\sim$50\% greater rate than prior methods and conditionally generates structures with desired space group properties. Our work demonstrates the potential of adapting post-training techniques from natural language processing to materials design, paving the way for targeted and efficient discovery of novel materials.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07150
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PLaID++: A Preference Aligned Language Model for Targeted Inorganic Materials Design
Xu, Andy
Desai, Rohan
Wang, Larry
Hope, Gabriel
Ritz, Ethan
Machine Learning
Materials Science
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising approach to improve correctness in LLMs, however, in many scientific problems, the objective is not necessarily to produce the correct answer, but instead to produce a diverse array of candidates which satisfy a set of constraints. We study this challenge in the context of materials generation. To this end, we introduce PLaID++, an LLM post-trained for stable and property-guided crystal generation. We find that performance hinges on our crystallographic representation and reward formulation. First, we introduce a compact, symmetry-informed Wyckoff text representation which improves computational efficiency and encourages generalization from physical priors. Second, we demonstrate that temperature scaling acts as an entropy regularizer which counteracts mode collapse and encourages exploration. By encoding symmetry constraints directly into text and guiding model outputs towards desirable chemical space, PLaID++ generates structures that are thermodynamically stable, unique, and novel at a $\sim$50\% greater rate than prior methods and conditionally generates structures with desired space group properties. Our work demonstrates the potential of adapting post-training techniques from natural language processing to materials design, paving the way for targeted and efficient discovery of novel materials.
title PLaID++: A Preference Aligned Language Model for Targeted Inorganic Materials Design
topic Machine Learning
Materials Science
url https://arxiv.org/abs/2509.07150