Structure-Aware Fill-in-the-Middle Pretraining for Code

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gong, Linyuan, Cheung, Alvin, Elhoushi, Mostafa, Wang, Sida
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915315519062016
author Gong, Linyuan
Cheung, Alvin
Elhoushi, Mostafa
Wang, Sida
author_facet Gong, Linyuan
Cheung, Alvin
Elhoushi, Mostafa
Wang, Sida
contents Fill-in-the-Middle (FIM) is a common pretraining method for code LLMs, where models complete code segments given surrounding context. However, existing LLMs treat code as plain text and mask random character spans. We propose and evaluate AST-FIM, a pretraining strategy that leverages Abstract Syntax Trees (ASTs) to mask complete syntactic structures at scale, ensuring coherent training examples better aligned with universal code structures and common code editing patterns such as blocks, expressions, or functions. To evaluate real-world fill-in-the-middle (FIM) programming tasks, we introduce Real-FIM-Eval, a benchmark derived from 30,000+ GitHub commits across 12 languages. On infilling tasks, experiments on 1B and 8B parameter models show that AST-FIM is particularly beneficial for real-world code editing as it outperforms standard random-character FIM by up to 5 pts on standard FIM benchmarks. Our code is publicly available at https://github.com/gonglinyuan/ast_fim.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00204
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structure-Aware Fill-in-the-Middle Pretraining for Code
Gong, Linyuan
Cheung, Alvin
Elhoushi, Mostafa
Wang, Sida
Computation and Language
Artificial Intelligence
Software Engineering
Fill-in-the-Middle (FIM) is a common pretraining method for code LLMs, where models complete code segments given surrounding context. However, existing LLMs treat code as plain text and mask random character spans. We propose and evaluate AST-FIM, a pretraining strategy that leverages Abstract Syntax Trees (ASTs) to mask complete syntactic structures at scale, ensuring coherent training examples better aligned with universal code structures and common code editing patterns such as blocks, expressions, or functions. To evaluate real-world fill-in-the-middle (FIM) programming tasks, we introduce Real-FIM-Eval, a benchmark derived from 30,000+ GitHub commits across 12 languages. On infilling tasks, experiments on 1B and 8B parameter models show that AST-FIM is particularly beneficial for real-world code editing as it outperforms standard random-character FIM by up to 5 pts on standard FIM benchmarks. Our code is publicly available at https://github.com/gonglinyuan/ast_fim.
title Structure-Aware Fill-in-the-Middle Pretraining for Code
topic Computation and Language
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2506.00204