Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ouédraogo, Wendkûuni C., Li, Yinghua, Dang, Xueqi, Zhou, Xin, Koyuncu, Anil, Klein, Jacques, Lo, David, Bissyandé, Tegawendé F.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908601156632576
author Ouédraogo, Wendkûuni C.
Li, Yinghua
Dang, Xueqi
Zhou, Xin
Koyuncu, Anil
Klein, Jacques
Lo, David
Bissyandé, Tegawendé F.
author_facet Ouédraogo, Wendkûuni C.
Li, Yinghua
Dang, Xueqi
Zhou, Xin
Koyuncu, Anil
Klein, Jacques
Lo, David
Bissyandé, Tegawendé F.
contents Large Language Models (LLMs) are increasingly used to refactor unit tests, improving readability and structure while preserving behavior. Evaluating such refactorings, however, remains difficult: metrics like CodeBLEU penalize beneficial renamings and edits, while semantic similarities overlook readability and modularity. We propose CTSES, a first step toward human-aligned evaluation of refactored tests. CTSES combines CodeBLEU, METEOR, and ROUGE-L into a composite score that balances semantics, lexical clarity, and structural alignment. Evaluated on 5,000+ refactorings from Defects4J and SF110 (GPT-4o and Mistral-Large), CTSES reduces false negatives and provides more interpretable signals than individual metrics. Our emerging results illustrate that CTSES offers a proof-of-concept for composite approaches, showing their promise in bridging automated metrics and developer judgments.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06767
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness
Ouédraogo, Wendkûuni C.
Li, Yinghua
Dang, Xueqi
Zhou, Xin
Koyuncu, Anil
Klein, Jacques
Lo, David
Bissyandé, Tegawendé F.
Software Engineering
Large Language Models (LLMs) are increasingly used to refactor unit tests, improving readability and structure while preserving behavior. Evaluating such refactorings, however, remains difficult: metrics like CodeBLEU penalize beneficial renamings and edits, while semantic similarities overlook readability and modularity. We propose CTSES, a first step toward human-aligned evaluation of refactored tests. CTSES combines CodeBLEU, METEOR, and ROUGE-L into a composite score that balances semantics, lexical clarity, and structural alignment. Evaluated on 5,000+ refactorings from Defects4J and SF110 (GPT-4o and Mistral-Large), CTSES reduces false negatives and provides more interpretable signals than individual metrics. Our emerging results illustrate that CTSES offers a proof-of-concept for composite approaches, showing their promise in bridging automated metrics and developer judgments.
title Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness
topic Software Engineering
url https://arxiv.org/abs/2506.06767