Incorporating Domain Knowledge into Materials Tokenization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oh, Yerim, Park, Jun-Hyung, Kim, Junho, Kim, SungHo, Lee, SangKeun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918057353412608
author Oh, Yerim
Park, Jun-Hyung
Kim, Junho
Kim, SungHo
Lee, SangKeun
author_facet Oh, Yerim
Park, Jun-Hyung
Kim, Junho
Kim, SungHo
Lee, SangKeun
contents While language models are increasingly utilized in materials science, typical models rely on frequency-centric tokenization methods originally developed for natural language processing. However, these methods frequently produce excessive fragmentation and semantic loss, failing to maintain the structural and semantic integrity of material concepts. To address this issue, we propose MATTER, a novel tokenization approach that integrates material knowledge into tokenization. Based on MatDetector trained on our materials knowledge base and a re-ranking method prioritizing material concepts in token merging, MATTER maintains the structural integrity of identified material concepts and prevents fragmentation during tokenization, ensuring their semantic meaning remains intact. The experimental results demonstrate that MATTER outperforms existing tokenization methods, achieving an average performance gain of $4\%$ and $2\%$ in the generation and classification tasks, respectively. These results underscore the importance of domain knowledge for tokenization strategies in scientific text processing. Our code is available at https://github.com/yerimoh/MATTER
format Preprint
id arxiv_https___arxiv_org_abs_2506_11115
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Incorporating Domain Knowledge into Materials Tokenization
Oh, Yerim
Park, Jun-Hyung
Kim, Junho
Kim, SungHo
Lee, SangKeun
Computation and Language
Artificial Intelligence
While language models are increasingly utilized in materials science, typical models rely on frequency-centric tokenization methods originally developed for natural language processing. However, these methods frequently produce excessive fragmentation and semantic loss, failing to maintain the structural and semantic integrity of material concepts. To address this issue, we propose MATTER, a novel tokenization approach that integrates material knowledge into tokenization. Based on MatDetector trained on our materials knowledge base and a re-ranking method prioritizing material concepts in token merging, MATTER maintains the structural integrity of identified material concepts and prevents fragmentation during tokenization, ensuring their semantic meaning remains intact. The experimental results demonstrate that MATTER outperforms existing tokenization methods, achieving an average performance gain of $4\%$ and $2\%$ in the generation and classification tasks, respectively. These results underscore the importance of domain knowledge for tokenization strategies in scientific text processing. Our code is available at https://github.com/yerimoh/MATTER
title Incorporating Domain Knowledge into Materials Tokenization
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.11115