Salvato in:
Dettagli Bibliografici
Autori principali: Nahri, Insaf, Pinquié, Romain, Véron, Philippe, Bus, Nicolas, Thorel, Mathieu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2508.13833
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909743417655296
author Nahri, Insaf
Pinquié, Romain
Véron, Philippe
Bus, Nicolas
Thorel, Mathieu
author_facet Nahri, Insaf
Pinquié, Romain
Véron, Philippe
Bus, Nicolas
Thorel, Mathieu
contents This study explores the integration of Building Information Modeling (BIM) with Natural Language Processing (NLP) to automate the extraction of requirements from unstructured French Building Technical Specification (BTS) documents within the construction industry. Employing Named Entity Recognition (NER) and Relation Extraction (RE) techniques, the study leverages the transformer-based model CamemBERT and applies transfer learning with the French language model Fr\_core\_news\_lg, both pre-trained on a large French corpus in the general domain. To benchmark these models, additional approaches ranging from rule-based to deep learning-based methods are developed. For RE, four different supervised models, including Random Forest, are implemented using a custom feature vector. A hand-crafted annotated dataset is used to compare the effectiveness of NER approaches and RE models. Results indicate that CamemBERT and Fr\_core\_news\_lg exhibited superior performance in NER, achieving F1-scores over 90\%, while Random Forest proved most effective in RE, with an F1 score above 80\%. The outcomes are intended to be represented as a knowledge graph in future work to further enhance automatic verification systems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13833
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Extracting Structured Requirements from Unstructured Building Technical Specifications for Building Information Modeling
Nahri, Insaf
Pinquié, Romain
Véron, Philippe
Bus, Nicolas
Thorel, Mathieu
Computation and Language
Artificial Intelligence
This study explores the integration of Building Information Modeling (BIM) with Natural Language Processing (NLP) to automate the extraction of requirements from unstructured French Building Technical Specification (BTS) documents within the construction industry. Employing Named Entity Recognition (NER) and Relation Extraction (RE) techniques, the study leverages the transformer-based model CamemBERT and applies transfer learning with the French language model Fr\_core\_news\_lg, both pre-trained on a large French corpus in the general domain. To benchmark these models, additional approaches ranging from rule-based to deep learning-based methods are developed. For RE, four different supervised models, including Random Forest, are implemented using a custom feature vector. A hand-crafted annotated dataset is used to compare the effectiveness of NER approaches and RE models. Results indicate that CamemBERT and Fr\_core\_news\_lg exhibited superior performance in NER, achieving F1-scores over 90\%, while Random Forest proved most effective in RE, with an F1 score above 80\%. The outcomes are intended to be represented as a knowledge graph in future work to further enhance automatic verification systems.
title Extracting Structured Requirements from Unstructured Building Technical Specifications for Building Information Modeling
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.13833