Saved in:
Bibliographic Details
Main Authors: Malyi, Max, Shek, Jonathan, McDonald, Alasdair, Biscaya, Andre
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.06813
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913023282642944
author Malyi, Max
Shek, Jonathan
McDonald, Alasdair
Biscaya, Andre
author_facet Malyi, Max
Shek, Jonathan
McDonald, Alasdair
Biscaya, Andre
contents Effective Operation and Maintenance (O&M) is critical to reducing the Levelised Cost of Energy (LCOE) from wind power, yet the unstructured, free-text nature of turbine maintenance logs presents a significant barrier to automated analysis. Our paper addresses this by presenting a novel and reproducible framework for benchmarking Large Language Models (LLMs) on the task of classifying these complex industrial records. To promote transparency and encourage further research, this framework has been made publicly available as an open-source tool. We systematically evaluate a diverse suite of state-of-the-art proprietary and open-source LLMs, providing a foundational assessment of their trade-offs in reliability, operational efficiency, and model calibration. Our results quantify a clear performance hierarchy, identifying top models that exhibit high alignment with a benchmark standard and trustworthy, well-calibrated confidence scores. We also demonstrate that classification performance is highly dependent on the task's semantic ambiguity, with all models showing higher consensus on objective component identification than on interpretive maintenance actions. Given that no model achieves perfect accuracy and that calibration varies dramatically, we conclude that the most effective and responsible near-term application is a Human-in-the-Loop system, where LLMs act as a powerful assistant to accelerate and standardise data labelling for human experts, thereby enhancing O&M data quality and downstream reliability analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06813
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Comparative Benchmark of Large Language Models for Labelling Wind Turbine Maintenance Logs
Malyi, Max
Shek, Jonathan
McDonald, Alasdair
Biscaya, Andre
Computation and Language
Effective Operation and Maintenance (O&M) is critical to reducing the Levelised Cost of Energy (LCOE) from wind power, yet the unstructured, free-text nature of turbine maintenance logs presents a significant barrier to automated analysis. Our paper addresses this by presenting a novel and reproducible framework for benchmarking Large Language Models (LLMs) on the task of classifying these complex industrial records. To promote transparency and encourage further research, this framework has been made publicly available as an open-source tool. We systematically evaluate a diverse suite of state-of-the-art proprietary and open-source LLMs, providing a foundational assessment of their trade-offs in reliability, operational efficiency, and model calibration. Our results quantify a clear performance hierarchy, identifying top models that exhibit high alignment with a benchmark standard and trustworthy, well-calibrated confidence scores. We also demonstrate that classification performance is highly dependent on the task's semantic ambiguity, with all models showing higher consensus on objective component identification than on interpretive maintenance actions. Given that no model achieves perfect accuracy and that calibration varies dramatically, we conclude that the most effective and responsible near-term application is a Human-in-the-Loop system, where LLMs act as a powerful assistant to accelerate and standardise data labelling for human experts, thereby enhancing O&M data quality and downstream reliability analysis.
title A Comparative Benchmark of Large Language Models for Labelling Wind Turbine Maintenance Logs
topic Computation and Language
url https://arxiv.org/abs/2509.06813