LLMORPH: Automated Metamorphic Testing of Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cho, Steven, Ruberto, Stefano, Terragni, Valerio
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912981368963072
author Cho, Steven
Ruberto, Stefano
Terragni, Valerio
author_facet Cho, Steven
Ruberto, Stefano
Terragni, Valerio
contents Automated testing is essential for evaluating and improving the reliability of Large Language Models (LLMs), yet the lack of automated oracles for verifying output correctness remains a key challenge. We present LLMORPH, an automated testing tool specifically designed for LLMs performing NLP tasks, which leverages Metamorphic Testing (MT) to uncover faulty behaviors without relying on human-labeled data. MT uses Metamorphic Relations (MRs) to generate follow-up inputs from source test input, enabling detection of inconsistencies in model outputs without the need of expensive labelled data. LLMORPH is aimed at researchers and developers who want to evaluate the robustness of LLM-based NLP systems. In this paper, we detail the design, implementation, and practical usage of LLMORPH, demonstrating how it can be easily extended to any LLM, NLP task, and set of MRs. In our evaluation, we applied 36 MRs across four NLP benchmarks, testing three state-of-the-art LLMs: GPT-4, LLAMA3, and HERMES 2. This produced over 561,000 test executions. Results demonstrate LLMORPH's effectiveness in automatically exposing inconsistencies.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23611
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLMORPH: Automated Metamorphic Testing of Large Language Models
Cho, Steven
Ruberto, Stefano
Terragni, Valerio
Software Engineering
Artificial Intelligence
Computation and Language
Machine Learning
D.2.5; I.2.7
Automated testing is essential for evaluating and improving the reliability of Large Language Models (LLMs), yet the lack of automated oracles for verifying output correctness remains a key challenge. We present LLMORPH, an automated testing tool specifically designed for LLMs performing NLP tasks, which leverages Metamorphic Testing (MT) to uncover faulty behaviors without relying on human-labeled data. MT uses Metamorphic Relations (MRs) to generate follow-up inputs from source test input, enabling detection of inconsistencies in model outputs without the need of expensive labelled data. LLMORPH is aimed at researchers and developers who want to evaluate the robustness of LLM-based NLP systems. In this paper, we detail the design, implementation, and practical usage of LLMORPH, demonstrating how it can be easily extended to any LLM, NLP task, and set of MRs. In our evaluation, we applied 36 MRs across four NLP benchmarks, testing three state-of-the-art LLMs: GPT-4, LLAMA3, and HERMES 2. This produced over 561,000 test executions. Results demonstrate LLMORPH's effectiveness in automatically exposing inconsistencies.
title LLMORPH: Automated Metamorphic Testing of Large Language Models
topic Software Engineering
Artificial Intelligence
Computation and Language
Machine Learning
D.2.5; I.2.7
url https://arxiv.org/abs/2603.23611