Scientific Reasoning: Assessment of Multimodal Generative LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dreyer, Florian, Kolos, Ekaterina, Matiash, Daria
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908253011574784
author Dreyer, Florian
Kolos, Ekaterina
Matiash, Daria
author_facet Dreyer, Florian
Kolos, Ekaterina
Matiash, Daria
contents Large language models (LLMs) can answer questions and reason about complex tasks, also from the scientific domain. We assess several multimodal LLMs (MLLMs) on ScienceQA and find that Gemini models show the highest accuracy with little context, and the highest textual similarity to human explanations with richer context. Adapter-tuning of smaller MLLMs did not lead to any reliable performance. Training from Gemini outputs consistently underperformed training from the original data.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01064
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scientific Reasoning: Assessment of Multimodal Generative LLMs
Dreyer, Florian
Kolos, Ekaterina
Matiash, Daria
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Large language models (LLMs) can answer questions and reason about complex tasks, also from the scientific domain. We assess several multimodal LLMs (MLLMs) on ScienceQA and find that Gemini models show the highest accuracy with little context, and the highest textual similarity to human explanations with richer context. Adapter-tuning of smaller MLLMs did not lead to any reliable performance. Training from Gemini outputs consistently underperformed training from the original data.
title Scientific Reasoning: Assessment of Multimodal Generative LLMs
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.01064