Rethinking Repetition Problems of LLMs in Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Yihong, Liu, Yuchen, Jiang, Xue, Jin, Zhi, Li, Ge
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909611301273600
author Dong, Yihong
Liu, Yuchen
Jiang, Xue
Jin, Zhi
Li, Ge
author_facet Dong, Yihong
Liu, Yuchen
Jiang, Xue
Jin, Zhi
Li, Ge
contents With the advent of neural language models, the performance of code generation has been significantly boosted. However, the problem of repetitions during the generation process continues to linger. Previous work has primarily focused on content repetition, which is merely a fraction of the broader repetition problem in code generation. A more prevalent and challenging problem is structural repetition. In structural repetition, the repeated code appears in various patterns but possesses a fixed structure, which can be inherently reflected in grammar. In this paper, we formally define structural repetition and propose an efficient decoding approach called RPG, which stands for Repetition Penalization based on Grammar, to alleviate the repetition problems in code generation for LLMs. Specifically, RPG first leverages grammar rules to identify repetition problems during code generation, and then strategically decays the likelihood of critical tokens that contribute to repetitions, thereby mitigating them in code generation. To facilitate this study, we construct a new dataset CodeRepetEval to comprehensively evaluate approaches for mitigating the repetition problems in code generation. Extensive experimental results demonstrate that RPG substantially outperforms the best-performing baselines on CodeRepetEval dataset as well as HumanEval and MBPP benchmarks, effectively reducing repetitions and enhancing the quality of generated code.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10402
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Repetition Problems of LLMs in Code Generation
Dong, Yihong
Liu, Yuchen
Jiang, Xue
Jin, Zhi
Li, Ge
Computation and Language
Artificial Intelligence
Machine Learning
Software Engineering
With the advent of neural language models, the performance of code generation has been significantly boosted. However, the problem of repetitions during the generation process continues to linger. Previous work has primarily focused on content repetition, which is merely a fraction of the broader repetition problem in code generation. A more prevalent and challenging problem is structural repetition. In structural repetition, the repeated code appears in various patterns but possesses a fixed structure, which can be inherently reflected in grammar. In this paper, we formally define structural repetition and propose an efficient decoding approach called RPG, which stands for Repetition Penalization based on Grammar, to alleviate the repetition problems in code generation for LLMs. Specifically, RPG first leverages grammar rules to identify repetition problems during code generation, and then strategically decays the likelihood of critical tokens that contribute to repetitions, thereby mitigating them in code generation. To facilitate this study, we construct a new dataset CodeRepetEval to comprehensively evaluate approaches for mitigating the repetition problems in code generation. Extensive experimental results demonstrate that RPG substantially outperforms the best-performing baselines on CodeRepetEval dataset as well as HumanEval and MBPP benchmarks, effectively reducing repetitions and enhancing the quality of generated code.
title Rethinking Repetition Problems of LLMs in Code Generation
topic Computation and Language
Artificial Intelligence
Machine Learning
Software Engineering
url https://arxiv.org/abs/2505.10402