Saved in:
Bibliographic Details
Main Authors: Yao, Junchi, Yang, Shu, Xu, Jianhua, Hu, Lijie, Li, Mengdi, Wang, Di
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.14218
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908586128441344
author Yao, Junchi
Yang, Shu
Xu, Jianhua
Hu, Lijie
Li, Mengdi
Wang, Di
author_facet Yao, Junchi
Yang, Shu
Xu, Jianhua
Hu, Lijie
Li, Mengdi
Wang, Di
contents Large language models (LLMs) have made remarkable progress in various domains, yet they often suffer from repetitive text generation, a phenomenon we refer to as the "Repeat Curse". While previous studies have proposed decoding strategies to mitigate repetition, the underlying mechanism behind this issue remains insufficiently explored. In this work, we investigate the root causes of repetition in LLMs through the lens of mechanistic interpretability. Inspired by recent advances in Sparse Autoencoders (SAEs), which enable monosemantic feature extraction, we propose a novel approach, "Duplicatus Charm", to induce and analyze the Repeat Curse. Our method systematically identifies "Repetition Features" -the key model activations responsible for generating repetitive outputs. First, we locate the layers most involved in repetition through logit analysis. Next, we extract and stimulate relevant features using SAE-based activation manipulation. To validate our approach, we construct a repetition dataset covering token and paragraph level repetitions and introduce an evaluation pipeline to quantify the influence of identified repetition features. Furthermore, by deactivating these features, we have effectively mitigated the Repeat Curse. The source code of our work is publicly available at: https://github.com/kaustpradalab/repeat-curse-llm
format Preprint
id arxiv_https___arxiv_org_abs_2504_14218
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding the Repeat Curse in Large Language Models from a Feature Perspective
Yao, Junchi
Yang, Shu
Xu, Jianhua
Hu, Lijie
Li, Mengdi
Wang, Di
Computation and Language
Large language models (LLMs) have made remarkable progress in various domains, yet they often suffer from repetitive text generation, a phenomenon we refer to as the "Repeat Curse". While previous studies have proposed decoding strategies to mitigate repetition, the underlying mechanism behind this issue remains insufficiently explored. In this work, we investigate the root causes of repetition in LLMs through the lens of mechanistic interpretability. Inspired by recent advances in Sparse Autoencoders (SAEs), which enable monosemantic feature extraction, we propose a novel approach, "Duplicatus Charm", to induce and analyze the Repeat Curse. Our method systematically identifies "Repetition Features" -the key model activations responsible for generating repetitive outputs. First, we locate the layers most involved in repetition through logit analysis. Next, we extract and stimulate relevant features using SAE-based activation manipulation. To validate our approach, we construct a repetition dataset covering token and paragraph level repetitions and introduce an evaluation pipeline to quantify the influence of identified repetition features. Furthermore, by deactivating these features, we have effectively mitigated the Repeat Curse. The source code of our work is publicly available at: https://github.com/kaustpradalab/repeat-curse-llm
title Understanding the Repeat Curse in Large Language Models from a Feature Perspective
topic Computation and Language
url https://arxiv.org/abs/2504.14218