Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Qinglong, Chen, Yuntian, Ma, Chao, Yang, Xiaokang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910010852769792
author Cao, Qinglong
Chen, Yuntian
Ma, Chao
Yang, Xiaokang
author_facet Cao, Qinglong
Chen, Yuntian
Ma, Chao
Yang, Xiaokang
contents Multimodal large language models (MLLMs) have shown remarkable capabilities in multimodal perception and understanding tasks. However, their effectiveness in specialized domains, such as remote sensing and medical imaging, remains limited. A natural approach to domain adaptation is to inject domain knowledge through textual instructions, prompts, or auxiliary captions. Surprisingly, we find that such input-level domain knowledge injection yields little to no improvement on scientific multimodal tasks, even when the domain knowledge is explicitly provided. This observation suggests that current MLLMs fail to internalize domain-specific priors through language alone, and that domain knowledge must be integrated at the optimization level. Motivated by this insight, we propose a reinforcement fine-tuning framework that incorporates domain knowledge directly into the learning objective. Instead of treating domain knowledge as descriptive information, we encode it as domain-informed constraints and reward signals, shaping the model's behavior in the output space. Extensive experiments across multiple datasets in remote sensing and medical domains consistently demonstrate good performance gains, achieving state-of-the-art results on multimodal domain tasks. Our results highlight the necessity of optimization-level domain knowledge integration and reveal a fundamental limitation of textual domain conditioning in current MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16419
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning
Cao, Qinglong
Chen, Yuntian
Ma, Chao
Yang, Xiaokang
Computation and Language
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) have shown remarkable capabilities in multimodal perception and understanding tasks. However, their effectiveness in specialized domains, such as remote sensing and medical imaging, remains limited. A natural approach to domain adaptation is to inject domain knowledge through textual instructions, prompts, or auxiliary captions. Surprisingly, we find that such input-level domain knowledge injection yields little to no improvement on scientific multimodal tasks, even when the domain knowledge is explicitly provided. This observation suggests that current MLLMs fail to internalize domain-specific priors through language alone, and that domain knowledge must be integrated at the optimization level. Motivated by this insight, we propose a reinforcement fine-tuning framework that incorporates domain knowledge directly into the learning objective. Instead of treating domain knowledge as descriptive information, we encode it as domain-informed constraints and reward signals, shaping the model's behavior in the output space. Extensive experiments across multiple datasets in remote sensing and medical domains consistently demonstrate good performance gains, achieving state-of-the-art results on multimodal domain tasks. Our results highlight the necessity of optimization-level domain knowledge integration and reveal a fundamental limitation of textual domain conditioning in current MLLMs.
title Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.16419