MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Durmaz, Ali Riza, Thomas, Akhil, Mishra, Lokesh, Murthy, Rachana Niranjan, Straub, Thomas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914906033356800
author Durmaz, Ali Riza
Thomas, Akhil
Mishra, Lokesh
Murthy, Rachana Niranjan
Straub, Thomas
author_facet Durmaz, Ali Riza
Thomas, Akhil
Mishra, Lokesh
Murthy, Rachana Niranjan
Straub, Thomas
contents While large language models learn sound statistical representations of the language and information therein, ontologies are symbolic knowledge representations that can complement the former ideally. Research at this critical intersection relies on datasets that intertwine ontologies and text corpora to enable training and comprehensive benchmarking of neurosymbolic models. We present the MaterioMiner dataset and the linked materials mechanics ontology where ontological concepts from the mechanics of materials domain are associated with textual entities within the literature corpus. Another distinctive feature of the dataset is its eminently fine-granular annotation. Specifically, 179 distinct classes are manually annotated by three raters within four publications, amounting to a total of 2191 entities that were annotated and curated. Conceptual work is presented for the symbolic representation of causal composition-process-microstructure-property relationships. We explore the annotation consistency between the three raters and perform fine-tuning of pre-trained models to showcase the feasibility of named-entity recognition model training. Reusing the dataset can foster training and benchmarking of materials language models, automated ontology construction, and knowledge graph generation from textual data.
format Preprint
id arxiv_https___arxiv_org_abs_2408_04661
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities
Durmaz, Ali Riza
Thomas, Akhil
Mishra, Lokesh
Murthy, Rachana Niranjan
Straub, Thomas
Computation and Language
Materials Science
While large language models learn sound statistical representations of the language and information therein, ontologies are symbolic knowledge representations that can complement the former ideally. Research at this critical intersection relies on datasets that intertwine ontologies and text corpora to enable training and comprehensive benchmarking of neurosymbolic models. We present the MaterioMiner dataset and the linked materials mechanics ontology where ontological concepts from the mechanics of materials domain are associated with textual entities within the literature corpus. Another distinctive feature of the dataset is its eminently fine-granular annotation. Specifically, 179 distinct classes are manually annotated by three raters within four publications, amounting to a total of 2191 entities that were annotated and curated. Conceptual work is presented for the symbolic representation of causal composition-process-microstructure-property relationships. We explore the annotation consistency between the three raters and perform fine-tuning of pre-trained models to showcase the feasibility of named-entity recognition model training. Reusing the dataset can foster training and benchmarking of materials language models, automated ontology construction, and knowledge graph generation from textual data.
title MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities
topic Computation and Language
Materials Science
url https://arxiv.org/abs/2408.04661