Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zou, Jiaru, Ban, Yikun, Li, Zihao, Qi, Yunzhe, Qiu, Ruizhong, Yang, Ling, He, Jingrui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912708214915072
author Zou, Jiaru
Ban, Yikun
Li, Zihao
Qi, Yunzhe
Qiu, Ruizhong
Yang, Ling
He, Jingrui
author_facet Zou, Jiaru
Ban, Yikun
Li, Zihao
Qi, Yunzhe
Qiu, Ruizhong
Yang, Ling
He, Jingrui
contents Large language models are typically adapted to downstream tasks through supervised fine-tuning on domain-specific data. While standard fine-tuning focuses on minimizing generation loss to optimize model parameters, we take a deeper step by retaining and leveraging the model's own learning signals, analogous to how human learners reflect on past mistakes to improve future performance. We first introduce the concept of Mistake Log to systematically track the model's learning behavior and recurring errors throughout fine-tuning. Treating the original transformer-based model as the Pilot, we correspondingly design a Copilot model to refine the Pilot's inference performance via logits rectification. We name the overall Pilot-Copilot framework the Transformer Copilot, which introduces (i) a novel Copilot model design, (ii) a joint training paradigm where the Copilot continuously learns from the evolving Mistake Log alongside the Pilot, and (iii) a fused inference paradigm where the Copilot rectifies the Pilot's logits for enhanced generation. We provide both theoretical and empirical analyses on our new learning framework. Experiments on 12 benchmarks spanning commonsense, arithmetic, and recommendation tasks demonstrate that Transformer Copilot consistently improves performance by up to 34.5%, while introducing marginal computational overhead to Pilot models and exhibiting strong scalability and transferability. Our code is released at https://github.com/jiaruzouu/TransformerCopilot.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning
Zou, Jiaru
Ban, Yikun
Li, Zihao
Qi, Yunzhe
Qiu, Ruizhong
Yang, Ling
He, Jingrui
Computation and Language
Artificial Intelligence
Machine Learning
Large language models are typically adapted to downstream tasks through supervised fine-tuning on domain-specific data. While standard fine-tuning focuses on minimizing generation loss to optimize model parameters, we take a deeper step by retaining and leveraging the model's own learning signals, analogous to how human learners reflect on past mistakes to improve future performance. We first introduce the concept of Mistake Log to systematically track the model's learning behavior and recurring errors throughout fine-tuning. Treating the original transformer-based model as the Pilot, we correspondingly design a Copilot model to refine the Pilot's inference performance via logits rectification. We name the overall Pilot-Copilot framework the Transformer Copilot, which introduces (i) a novel Copilot model design, (ii) a joint training paradigm where the Copilot continuously learns from the evolving Mistake Log alongside the Pilot, and (iii) a fused inference paradigm where the Copilot rectifies the Pilot's logits for enhanced generation. We provide both theoretical and empirical analyses on our new learning framework. Experiments on 12 benchmarks spanning commonsense, arithmetic, and recommendation tasks demonstrate that Transformer Copilot consistently improves performance by up to 34.5%, while introducing marginal computational overhead to Pilot models and exhibiting strong scalability and transferability. Our code is released at https://github.com/jiaruzouu/TransformerCopilot.
title Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.16270