Enhancing Code LLM Training with Programmer Attention

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Yifan, Huang, Chen, Karas, Zachary, Nguyen, Dung Thuy, Leach, Kevin, Huang, Yu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915242934534144
author Zhang, Yifan
Huang, Chen
Karas, Zachary
Nguyen, Dung Thuy
Leach, Kevin
Huang, Yu
author_facet Zhang, Yifan
Huang, Chen
Karas, Zachary
Nguyen, Dung Thuy
Leach, Kevin
Huang, Yu
contents Human attention provides valuable yet underexploited signals for code LLM training, offering a perspective beyond purely machine-driven attention. Despite the complexity and cost of collecting eye-tracking data, there has also been limited progress in systematically using these signals for code LLM training. To address both issues, we propose a cohesive pipeline spanning augmentation and reward-based fine-tuning. Specifically, we introduce (1) an eye-tracking path augmentation method to expand programmer attention datasets, (2) a pattern abstraction step that refines raw fixations into learnable attention motifs, and (3) a reward-guided strategy for integrating these insights directly into a CodeT5 supervised fine-tuning process. Our experiments yield +7.16 in CodeBLEU on the CodeXGlue benchmark for code summarization, underscoring how uniting human and machine attention can boost code intelligence. We hope this work encourages broader exploration of human-centric methods in next-generation AI4SE.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14936
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Code LLM Training with Programmer Attention
Zhang, Yifan
Huang, Chen
Karas, Zachary
Nguyen, Dung Thuy
Leach, Kevin
Huang, Yu
Software Engineering
Human-Computer Interaction
Machine Learning
Human attention provides valuable yet underexploited signals for code LLM training, offering a perspective beyond purely machine-driven attention. Despite the complexity and cost of collecting eye-tracking data, there has also been limited progress in systematically using these signals for code LLM training. To address both issues, we propose a cohesive pipeline spanning augmentation and reward-based fine-tuning. Specifically, we introduce (1) an eye-tracking path augmentation method to expand programmer attention datasets, (2) a pattern abstraction step that refines raw fixations into learnable attention motifs, and (3) a reward-guided strategy for integrating these insights directly into a CodeT5 supervised fine-tuning process. Our experiments yield +7.16 in CodeBLEU on the CodeXGlue benchmark for code summarization, underscoring how uniting human and machine attention can boost code intelligence. We hope this work encourages broader exploration of human-centric methods in next-generation AI4SE.
title Enhancing Code LLM Training with Programmer Attention
topic Software Engineering
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2503.14936