Rethinking Cognitive Complexity for Unit Tests: Toward a Readability-Aware Metric Grounded in Developer Perception

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ouédraogo, Wendkûuni C., Li, Yinghua, Dang, Xueqi, Zhou, Xin, Koyuncu, Anil, Klein, Jacques, Lo, David, Bissyandé, Tegawendé F.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911118334623744
author Ouédraogo, Wendkûuni C.
Li, Yinghua
Dang, Xueqi
Zhou, Xin
Koyuncu, Anil
Klein, Jacques
Lo, David
Bissyandé, Tegawendé F.
author_facet Ouédraogo, Wendkûuni C.
Li, Yinghua
Dang, Xueqi
Zhou, Xin
Koyuncu, Anil
Klein, Jacques
Lo, David
Bissyandé, Tegawendé F.
contents Automatically generated unit tests-from search-based tools like EvoSuite or LLMs-vary significantly in structure and readability. Yet most evaluations rely on metrics like Cyclomatic Complexity and Cognitive Complexity, designed for functional code rather than test code. Recent studies have shown that SonarSource's Cognitive Complexity metric assigns near-zero scores to LLM-generated tests, yet its behavior on EvoSuite-generated tests and its applicability to test-specific code structures remain unexplored. We introduce CCTR, a Test-Aware Cognitive Complexity metric tailored for unit tests. CCTR integrates structural and semantic features like assertion density, annotation roles, and test composition patterns-dimensions ignored by traditional complexity models but critical for understanding test code. We evaluate 15,750 test suites generated by EvoSuite, GPT-4o, and Mistral Large-1024 across 350 classes from Defects4J and SF110. Results show CCTR effectively discriminates between structured and fragmented test suites, producing interpretable scores that better reflect developer-perceived effort. By bridging structural analysis and test readability, CCTR provides a foundation for more reliable evaluation and improvement of generated tests. We publicly release all data, prompts, and evaluation scripts to support replication.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06764
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Cognitive Complexity for Unit Tests: Toward a Readability-Aware Metric Grounded in Developer Perception
Ouédraogo, Wendkûuni C.
Li, Yinghua
Dang, Xueqi
Zhou, Xin
Koyuncu, Anil
Klein, Jacques
Lo, David
Bissyandé, Tegawendé F.
Software Engineering
Automatically generated unit tests-from search-based tools like EvoSuite or LLMs-vary significantly in structure and readability. Yet most evaluations rely on metrics like Cyclomatic Complexity and Cognitive Complexity, designed for functional code rather than test code. Recent studies have shown that SonarSource's Cognitive Complexity metric assigns near-zero scores to LLM-generated tests, yet its behavior on EvoSuite-generated tests and its applicability to test-specific code structures remain unexplored. We introduce CCTR, a Test-Aware Cognitive Complexity metric tailored for unit tests. CCTR integrates structural and semantic features like assertion density, annotation roles, and test composition patterns-dimensions ignored by traditional complexity models but critical for understanding test code. We evaluate 15,750 test suites generated by EvoSuite, GPT-4o, and Mistral Large-1024 across 350 classes from Defects4J and SF110. Results show CCTR effectively discriminates between structured and fragmented test suites, producing interpretable scores that better reflect developer-perceived effort. By bridging structural analysis and test readability, CCTR provides a foundation for more reliable evaluation and improvement of generated tests. We publicly release all data, prompts, and evaluation scripts to support replication.
title Rethinking Cognitive Complexity for Unit Tests: Toward a Readability-Aware Metric Grounded in Developer Perception
topic Software Engineering
url https://arxiv.org/abs/2506.06764