Empowering Lightweight MLLMs with Reasoning via Long CoT SFT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ou, Linyu, Yin, YuYang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912637417160704
author Ou, Linyu
Yin, YuYang
author_facet Ou, Linyu
Yin, YuYang
contents While Reinforcement Learning with Verifiable Rewards has enhanced the reasoning of large-scale language models (LLMs), its efficacy for lightweight multimodal language models (MLLMs) with fewer than seven billion parameters remains underexplored. This paper investigates the role of long Chain-of-Thought (long CoT) data in enhancing the reasoning abilities of such MLLMs. Our findings demonstrate that Supervised Fine-Tuning (SFT) with long CoT data significantly improves MLLM reasoning. Furthermore, we observe that after this initial SFT phase, MLLMs can achieve additional performance gains through a subsequent RL stage. We conclude that a SFT stage with long CoT data is a critical prerequisite for developing the reasoning capabilities of lightweight MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03321
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
Ou, Linyu
Yin, YuYang
Computer Vision and Pattern Recognition
While Reinforcement Learning with Verifiable Rewards has enhanced the reasoning of large-scale language models (LLMs), its efficacy for lightweight multimodal language models (MLLMs) with fewer than seven billion parameters remains underexplored. This paper investigates the role of long Chain-of-Thought (long CoT) data in enhancing the reasoning abilities of such MLLMs. Our findings demonstrate that Supervised Fine-Tuning (SFT) with long CoT data significantly improves MLLM reasoning. Furthermore, we observe that after this initial SFT phase, MLLMs can achieve additional performance gains through a subsequent RL stage. We conclude that a SFT stage with long CoT data is a critical prerequisite for developing the reasoning capabilities of lightweight MLLMs.
title Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.03321