Scientific Reasoning: Assessment of Multimodal Generative LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908253011574784 |
|---|---|
| author | Dreyer, Florian Kolos, Ekaterina Matiash, Daria |
| author_facet | Dreyer, Florian Kolos, Ekaterina Matiash, Daria |
| contents | Large language models (LLMs) can answer questions and reason about complex tasks, also from the scientific domain. We assess several multimodal LLMs (MLLMs) on ScienceQA and find that Gemini models show the highest accuracy with little context, and the highest textual similarity to human explanations with richer context. Adapter-tuning of smaller MLLMs did not lead to any reliable performance. Training from Gemini outputs consistently underperformed training from the original data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_01064 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Scientific Reasoning: Assessment of Multimodal Generative LLMs Dreyer, Florian Kolos, Ekaterina Matiash, Daria Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Large language models (LLMs) can answer questions and reason about complex tasks, also from the scientific domain. We assess several multimodal LLMs (MLLMs) on ScienceQA and find that Gemini models show the highest accuracy with little context, and the highest textual similarity to human explanations with richer context. Adapter-tuning of smaller MLLMs did not lead to any reliable performance. Training from Gemini outputs consistently underperformed training from the original data. |
| title | Scientific Reasoning: Assessment of Multimodal Generative LLMs |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.01064 |