Evaluating ChatGPT's Performance in Classifying Pneumonia from Chest X-Ray Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Prahallad, Pragna, Prahallad, Pranathi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915576118509568
author Prahallad, Pragna
Prahallad, Pranathi
author_facet Prahallad, Pragna
Prahallad, Pranathi
contents In this study, we evaluate the ability of OpenAI's gpt-4o model to classify chest X-ray images as either NORMAL or PNEUMONIA in a zero-shot setting, without any prior fine-tuning. A balanced test set of 400 images (200 from each class) was used to assess performance across four distinct prompt designs, ranging from minimal instructions to detailed, reasoning-based prompts. The results indicate that concise, feature-focused prompts achieved the highest classification accuracy of 74\%, whereas reasoning-oriented prompts resulted in lower performance. These findings highlight that while ChatGPT exhibits emerging potential for medical image interpretation, its diagnostic reliability remains limited. Continued advances in visual reasoning and domain-specific adaptation are required before such models can be safely applied in clinical practice.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating ChatGPT's Performance in Classifying Pneumonia from Chest X-Ray Images
Prahallad, Pragna
Prahallad, Pranathi
Computer Vision and Pattern Recognition
Artificial Intelligence
In this study, we evaluate the ability of OpenAI's gpt-4o model to classify chest X-ray images as either NORMAL or PNEUMONIA in a zero-shot setting, without any prior fine-tuning. A balanced test set of 400 images (200 from each class) was used to assess performance across four distinct prompt designs, ranging from minimal instructions to detailed, reasoning-based prompts. The results indicate that concise, feature-focused prompts achieved the highest classification accuracy of 74\%, whereas reasoning-oriented prompts resulted in lower performance. These findings highlight that while ChatGPT exhibits emerging potential for medical image interpretation, its diagnostic reliability remains limited. Continued advances in visual reasoning and domain-specific adaptation are required before such models can be safely applied in clinical practice.
title Evaluating ChatGPT's Performance in Classifying Pneumonia from Chest X-Ray Images
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.21839