ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiaqi, Dong, Xinyi, Liu, Yang, Yang, Zhizhuo, Wang, Quansen, Wang, Xiaobo, Zhu, SongChun, Jia, Zixia, Zheng, Zilong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909619844022272
author Li, Jiaqi
Dong, Xinyi
Liu, Yang
Yang, Zhizhuo
Wang, Quansen
Wang, Xiaobo
Zhu, SongChun
Jia, Zixia
Zheng, Zilong
author_facet Li, Jiaqi
Dong, Xinyi
Liu, Yang
Yang, Zhizhuo
Wang, Quansen
Wang, Xiaobo
Zhu, SongChun
Jia, Zixia
Zheng, Zilong
contents We present a novel pipeline, ReflectEvo, to demonstrate that small language models (SLMs) can enhance meta introspection through reflection learning. This process iteratively generates self-reflection for self-training, fostering a continuous and self-evolving process. Leveraging this pipeline, we construct ReflectEvo-460k, a large-scale, comprehensive, self-generated reflection dataset with broadened instructions and diverse multi-domain tasks. Building upon this dataset, we demonstrate the effectiveness of reflection learning to improve SLMs' reasoning abilities using SFT and DPO with remarkable performance, substantially boosting Llama-3 from 52.4% to 71.2% and Mistral from 44.4% to 71.1%. It validates that ReflectEvo can rival or even surpass the reasoning capability of the three prominent open-sourced models on BIG-bench without distillation from superior models or fine-grained human annotation. We further conduct a deeper analysis of the high quality of self-generated reflections and their impact on error localization and correction. Our work highlights the potential of continuously enhancing the reasoning performance of SLMs through iterative reflection learning in the long run.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16475
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
Li, Jiaqi
Dong, Xinyi
Liu, Yang
Yang, Zhizhuo
Wang, Quansen
Wang, Xiaobo
Zhu, SongChun
Jia, Zixia
Zheng, Zilong
Artificial Intelligence
We present a novel pipeline, ReflectEvo, to demonstrate that small language models (SLMs) can enhance meta introspection through reflection learning. This process iteratively generates self-reflection for self-training, fostering a continuous and self-evolving process. Leveraging this pipeline, we construct ReflectEvo-460k, a large-scale, comprehensive, self-generated reflection dataset with broadened instructions and diverse multi-domain tasks. Building upon this dataset, we demonstrate the effectiveness of reflection learning to improve SLMs' reasoning abilities using SFT and DPO with remarkable performance, substantially boosting Llama-3 from 52.4% to 71.2% and Mistral from 44.4% to 71.1%. It validates that ReflectEvo can rival or even surpass the reasoning capability of the three prominent open-sourced models on BIG-bench without distillation from superior models or fine-grained human annotation. We further conduct a deeper analysis of the high quality of self-generated reflections and their impact on error localization and correction. Our work highlights the potential of continuously enhancing the reasoning performance of SLMs through iterative reflection learning in the long run.
title ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
topic Artificial Intelligence
url https://arxiv.org/abs/2505.16475