Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908601156632576 |
|---|---|
| author | Ouédraogo, Wendkûuni C. Li, Yinghua Dang, Xueqi Zhou, Xin Koyuncu, Anil Klein, Jacques Lo, David Bissyandé, Tegawendé F. |
| author_facet | Ouédraogo, Wendkûuni C. Li, Yinghua Dang, Xueqi Zhou, Xin Koyuncu, Anil Klein, Jacques Lo, David Bissyandé, Tegawendé F. |
| contents | Large Language Models (LLMs) are increasingly used to refactor unit tests, improving readability and structure while preserving behavior. Evaluating such refactorings, however, remains difficult: metrics like CodeBLEU penalize beneficial renamings and edits, while semantic similarities overlook readability and modularity. We propose CTSES, a first step toward human-aligned evaluation of refactored tests. CTSES combines CodeBLEU, METEOR, and ROUGE-L into a composite score that balances semantics, lexical clarity, and structural alignment. Evaluated on 5,000+ refactorings from Defects4J and SF110 (GPT-4o and Mistral-Large), CTSES reduces false negatives and provides more interpretable signals than individual metrics. Our emerging results illustrate that CTSES offers a proof-of-concept for composite approaches, showing their promise in bridging automated metrics and developer judgments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_06767 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness Ouédraogo, Wendkûuni C. Li, Yinghua Dang, Xueqi Zhou, Xin Koyuncu, Anil Klein, Jacques Lo, David Bissyandé, Tegawendé F. Software Engineering Large Language Models (LLMs) are increasingly used to refactor unit tests, improving readability and structure while preserving behavior. Evaluating such refactorings, however, remains difficult: metrics like CodeBLEU penalize beneficial renamings and edits, while semantic similarities overlook readability and modularity. We propose CTSES, a first step toward human-aligned evaluation of refactored tests. CTSES combines CodeBLEU, METEOR, and ROUGE-L into a composite score that balances semantics, lexical clarity, and structural alignment. Evaluated on 5,000+ refactorings from Defects4J and SF110 (GPT-4o and Mistral-Large), CTSES reduces false negatives and provides more interpretable signals than individual metrics. Our emerging results illustrate that CTSES offers a proof-of-concept for composite approaches, showing their promise in bridging automated metrics and developer judgments. |
| title | Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness |
| topic | Software Engineering |
| url | https://arxiv.org/abs/2506.06767 |