LLMBoost: Make Large Language Models Stronger with Boosting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zehao, Ai, Tianxiang, Li, Yifei, Li, Gongxun, Wei, Yuyang, Zhou, Wang, Li, Guanghui, Yu, Bin, Chen, Zhijun, Sun, Hailong, Zhuang, Fuzhen, Li, Jianxin, Wang, Deqing, Ban, Yikun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909976713232384
author Chen, Zehao
Ai, Tianxiang
Li, Yifei
Li, Gongxun
Wei, Yuyang
Zhou, Wang
Li, Guanghui
Yu, Bin
Chen, Zhijun
Sun, Hailong
Zhuang, Fuzhen
Li, Jianxin
Wang, Deqing
Ban, Yikun
author_facet Chen, Zehao
Ai, Tianxiang
Li, Yifei
Li, Gongxun
Wei, Yuyang
Zhou, Wang
Li, Guanghui
Yu, Bin
Chen, Zhijun
Sun, Hailong
Zhuang, Fuzhen
Li, Jianxin
Wang, Deqing
Ban, Yikun
contents Ensemble learning of LLMs has emerged as a promising alternative to enhance performance, but existing approaches typically treat models as black boxes, combining the inputs or final outputs while overlooking the rich internal representations and interactions across models.In this work, we introduce LLMBoost, a novel ensemble fine-tuning framework that breaks this barrier by explicitly leveraging intermediate states of LLMs. Inspired by the boosting paradigm, LLMBoost incorporates three key innovations. First, a cross-model attention mechanism enables successor models to access and fuse hidden states from predecessors, facilitating hierarchical error correction and knowledge transfer. Second, a chain training paradigm progressively fine-tunes connected models with an error-suppression objective, ensuring that each model rectifies the mispredictions of its predecessor with minimal additional computation. Third, a near-parallel inference paradigm design pipelines hidden states across models layer by layer, achieving inference efficiency approaching single-model decoding. We further establish the theoretical foundations of LLMBoost, proving that sequential integration guarantees monotonic improvements under bounded correction assumptions. Extensive experiments on commonsense reasoning and arithmetic reasoning tasks demonstrate that LLMBoost consistently boosts accuracy while reducing inference latency.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22309
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMBoost: Make Large Language Models Stronger with Boosting
Chen, Zehao
Ai, Tianxiang
Li, Yifei
Li, Gongxun
Wei, Yuyang
Zhou, Wang
Li, Guanghui
Yu, Bin
Chen, Zhijun
Sun, Hailong
Zhuang, Fuzhen
Li, Jianxin
Wang, Deqing
Ban, Yikun
Machine Learning
Artificial Intelligence
Ensemble learning of LLMs has emerged as a promising alternative to enhance performance, but existing approaches typically treat models as black boxes, combining the inputs or final outputs while overlooking the rich internal representations and interactions across models.In this work, we introduce LLMBoost, a novel ensemble fine-tuning framework that breaks this barrier by explicitly leveraging intermediate states of LLMs. Inspired by the boosting paradigm, LLMBoost incorporates three key innovations. First, a cross-model attention mechanism enables successor models to access and fuse hidden states from predecessors, facilitating hierarchical error correction and knowledge transfer. Second, a chain training paradigm progressively fine-tunes connected models with an error-suppression objective, ensuring that each model rectifies the mispredictions of its predecessor with minimal additional computation. Third, a near-parallel inference paradigm design pipelines hidden states across models layer by layer, achieving inference efficiency approaching single-model decoding. We further establish the theoretical foundations of LLMBoost, proving that sequential integration guarantees monotonic improvements under bounded correction assumptions. Extensive experiments on commonsense reasoning and arithmetic reasoning tasks demonstrate that LLMBoost consistently boosts accuracy while reducing inference latency.
title LLMBoost: Make Large Language Models Stronger with Boosting
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.22309