MUCOCO: Automated Consistency Testing of Code LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chou, Chua Jin, Lwin, Khant That, Soremekun, Ezekiel
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913050891649024
author Chou, Chua Jin
Lwin, Khant That
Soremekun, Ezekiel
author_facet Chou, Chua Jin
Lwin, Khant That
Soremekun, Ezekiel
contents Code LLMs often portray inconsistent program behaviors. Developers typically employ benchmarks to assess Code LLMs, but most benchmarks are hand-crafted, static and do not target consistency property. In this work, we pose the scientific question: how can we automatically discover inconsistent program behaviors in Code LLMs? To address this challenge, we propose an automated consistency testing method, called MUCOCO, which employs semantic-preserving mutation analysis to expose inconsistent behaviors in code LLMs. Given a coding query, MUCOCO automatically transforms its program into semantically equivalent programs (aka mutants) and detects inconsistencies between the mutants and the original program (e.g., different output or test failure). We evaluate MUCOCO using four (4) coding tasks and seven (7) LLMs. Results show that MUCOCO is effective in exposing inconsistency and outperforms the closest baseline (TURBULENCE). About one in seven (15%) inputs generated by MUCOCO exposed inconsistencies. Our work motivates the need to test Code LLMs for consistency property
format Preprint
id arxiv_https___arxiv_org_abs_2604_19086
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MUCOCO: Automated Consistency Testing of Code LLMs
Chou, Chua Jin
Lwin, Khant That
Soremekun, Ezekiel
Software Engineering
Code LLMs often portray inconsistent program behaviors. Developers typically employ benchmarks to assess Code LLMs, but most benchmarks are hand-crafted, static and do not target consistency property. In this work, we pose the scientific question: how can we automatically discover inconsistent program behaviors in Code LLMs? To address this challenge, we propose an automated consistency testing method, called MUCOCO, which employs semantic-preserving mutation analysis to expose inconsistent behaviors in code LLMs. Given a coding query, MUCOCO automatically transforms its program into semantically equivalent programs (aka mutants) and detects inconsistencies between the mutants and the original program (e.g., different output or test failure). We evaluate MUCOCO using four (4) coding tasks and seven (7) LLMs. Results show that MUCOCO is effective in exposing inconsistency and outperforms the closest baseline (TURBULENCE). About one in seven (15%) inputs generated by MUCOCO exposed inconsistencies. Our work motivates the need to test Code LLMs for consistency property
title MUCOCO: Automated Consistency Testing of Code LLMs
topic Software Engineering
url https://arxiv.org/abs/2604.19086