The Impact of Image Resolution on Biomedical Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Liangyu, Burgess, James, Nirschl, Jeffrey J, Zohar, Orr, Yeung-Levy, Serena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912662672113664
author Chen, Liangyu
Burgess, James
Nirschl, Jeffrey J
Zohar, Orr
Yeung-Levy, Serena
author_facet Chen, Liangyu
Burgess, James
Nirschl, Jeffrey J
Zohar, Orr
Yeung-Levy, Serena
contents Imaging technologies are fundamental to biomedical research and modern medicine, requiring analysis of high-resolution images across various modalities. While multimodal large language models (MLLMs) show promise for biomedical image analysis, most are designed for low-resolution images from general-purpose datasets, risking critical information loss. We investigate how image resolution affects MLLM performance in biomedical applications and demonstrate that: (1) native-resolution training and inference significantly improve performance across multiple tasks, (2) misalignment between training and inference resolutions severely degrades performance, and (3) mixed-resolution training effectively mitigates misalignment and balances computational constraints with performance requirements. Based on these findings, we recommend prioritizing native-resolution inference and mixed-resolution datasets to optimize biomedical MLLMs for transformative impact in scientific research and clinical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Impact of Image Resolution on Biomedical Multimodal Large Language Models
Chen, Liangyu
Burgess, James
Nirschl, Jeffrey J
Zohar, Orr
Yeung-Levy, Serena
Computer Vision and Pattern Recognition
Computation and Language
Imaging technologies are fundamental to biomedical research and modern medicine, requiring analysis of high-resolution images across various modalities. While multimodal large language models (MLLMs) show promise for biomedical image analysis, most are designed for low-resolution images from general-purpose datasets, risking critical information loss. We investigate how image resolution affects MLLM performance in biomedical applications and demonstrate that: (1) native-resolution training and inference significantly improve performance across multiple tasks, (2) misalignment between training and inference resolutions severely degrades performance, and (3) mixed-resolution training effectively mitigates misalignment and balances computational constraints with performance requirements. Based on these findings, we recommend prioritizing native-resolution inference and mixed-resolution datasets to optimize biomedical MLLMs for transformative impact in scientific research and clinical applications.
title The Impact of Image Resolution on Biomedical Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.18304