Success is in the Details: Evaluate and Enhance Details Sensitivity of Code LLMs through Counterfactuals

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Luo, Xianzhen, Zhu, Qingfu, Zhang, Zhiming, Xu, Mingzheng, Cheng, Tianhao, Wang, Yixuan, Chu, Zheng, Xuyang, Shijie, Ma, Zhiyuan, Fan, YuanTao, Che, Wanxiang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910957847969792
author Luo, Xianzhen
Zhu, Qingfu
Zhang, Zhiming
Xu, Mingzheng
Cheng, Tianhao
Wang, Yixuan
Chu, Zheng
Xuyang, Shijie
Ma, Zhiyuan
Fan, YuanTao
Che, Wanxiang
author_facet Luo, Xianzhen
Zhu, Qingfu
Zhang, Zhiming
Xu, Mingzheng
Cheng, Tianhao
Wang, Yixuan
Chu, Zheng
Xuyang, Shijie
Ma, Zhiyuan
Fan, YuanTao
Che, Wanxiang
contents Code Sensitivity refers to the ability of Code LLMs to recognize and respond to details changes in problem descriptions. While current code benchmarks and instruction data focus on difficulty and diversity, sensitivity is overlooked. We first introduce the CTF-Code benchmark, constructed using counterfactual perturbations, minimizing input changes while maximizing output changes. The evaluation shows that many LLMs have a more than 10\% performance drop compared to the original problems. To fully utilize sensitivity, CTF-Instruct, an incremental instruction fine-tuning framework, extends on existing data and uses a selection mechanism to meet the three dimensions of difficulty, diversity, and sensitivity. Experiments show that LLMs fine-tuned with CTF-Instruct data achieve over a 2\% improvement on CTF-Code, and more than a 10\% performance boost on LiveCodeBench, validating the feasibility of enhancing LLMs' sensitivity to improve performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14597
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Success is in the Details: Evaluate and Enhance Details Sensitivity of Code LLMs through Counterfactuals
Luo, Xianzhen
Zhu, Qingfu
Zhang, Zhiming
Xu, Mingzheng
Cheng, Tianhao
Wang, Yixuan
Chu, Zheng
Xuyang, Shijie
Ma, Zhiyuan
Fan, YuanTao
Che, Wanxiang
Computation and Language
Code Sensitivity refers to the ability of Code LLMs to recognize and respond to details changes in problem descriptions. While current code benchmarks and instruction data focus on difficulty and diversity, sensitivity is overlooked. We first introduce the CTF-Code benchmark, constructed using counterfactual perturbations, minimizing input changes while maximizing output changes. The evaluation shows that many LLMs have a more than 10\% performance drop compared to the original problems. To fully utilize sensitivity, CTF-Instruct, an incremental instruction fine-tuning framework, extends on existing data and uses a selection mechanism to meet the three dimensions of difficulty, diversity, and sensitivity. Experiments show that LLMs fine-tuned with CTF-Instruct data achieve over a 2\% improvement on CTF-Code, and more than a 10\% performance boost on LiveCodeBench, validating the feasibility of enhancing LLMs' sensitivity to improve performance.
title Success is in the Details: Evaluate and Enhance Details Sensitivity of Code LLMs through Counterfactuals
topic Computation and Language
url https://arxiv.org/abs/2505.14597