From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MARKERGEN

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yuan, Peiwen, Tan, Chuyi, Feng, Shaoxiong, Li, Yiwei, Wang, Xinglin, Zhang, Yueqi, Shi, Jiayi, Pan, Boyuan, Hu, Yao, Li, Kan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910994010210304
author Yuan, Peiwen
Tan, Chuyi
Feng, Shaoxiong
Li, Yiwei
Wang, Xinglin
Zhang, Yueqi
Shi, Jiayi
Pan, Boyuan
Hu, Yao
Li, Kan
author_facet Yuan, Peiwen
Tan, Chuyi
Feng, Shaoxiong
Li, Yiwei
Wang, Xinglin
Zhang, Yueqi
Shi, Jiayi
Pan, Boyuan
Hu, Yao
Li, Kan
contents Despite the rapid progress of large language models (LLMs), their length-controllable text generation (LCTG) ability remains below expectations, posing a major limitation for practical applications. Existing methods mainly focus on end-to-end training to reinforce adherence to length constraints. However, the lack of decomposition and targeted enhancement of LCTG sub-abilities restricts further progress. To bridge this gap, we conduct a bottom-up decomposition of LCTG sub-abilities with human patterns as reference and perform a detailed error analysis. On this basis, we propose MarkerGen, a simple-yet-effective plug-and-play approach that:(1) mitigates LLM fundamental deficiencies via external tool integration;(2) conducts explicit length modeling with dynamically inserted markers;(3) employs a three-stage generation scheme to better align length constraints while maintaining content quality. Comprehensive experiments demonstrate that MarkerGen significantly improves LCTG across various settings, exhibiting outstanding effectiveness and generalizability.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13544
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MARKERGEN
Yuan, Peiwen
Tan, Chuyi
Feng, Shaoxiong
Li, Yiwei
Wang, Xinglin
Zhang, Yueqi
Shi, Jiayi
Pan, Boyuan
Hu, Yao
Li, Kan
Computation and Language
Artificial Intelligence
Despite the rapid progress of large language models (LLMs), their length-controllable text generation (LCTG) ability remains below expectations, posing a major limitation for practical applications. Existing methods mainly focus on end-to-end training to reinforce adherence to length constraints. However, the lack of decomposition and targeted enhancement of LCTG sub-abilities restricts further progress. To bridge this gap, we conduct a bottom-up decomposition of LCTG sub-abilities with human patterns as reference and perform a detailed error analysis. On this basis, we propose MarkerGen, a simple-yet-effective plug-and-play approach that:(1) mitigates LLM fundamental deficiencies via external tool integration;(2) conducts explicit length modeling with dynamically inserted markers;(3) employs a three-stage generation scheme to better align length constraints while maintaining content quality. Comprehensive experiments demonstrate that MarkerGen significantly improves LCTG across various settings, exhibiting outstanding effectiveness and generalizability.
title From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MARKERGEN
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.13544