FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Khan, Faizan Farooq, Radwan, Yousef, Abdelrahman, Eslam, Felemban, Abdulwahab, Mir, Aymen, Michiels, Nico K., Temple, Andrew J., Berumen, Michael L., Elhoseiny, Mohamed
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915523970727936
author Khan, Faizan Farooq
Radwan, Yousef
Abdelrahman, Eslam
Felemban, Abdulwahab
Mir, Aymen
Michiels, Nico K.
Temple, Andrew J.
Berumen, Michael L.
Elhoseiny, Mohamed
author_facet Khan, Faizan Farooq
Radwan, Yousef
Abdelrahman, Eslam
Felemban, Abdulwahab
Mir, Aymen
Michiels, Nico K.
Temple, Andrew J.
Berumen, Michael L.
Elhoseiny, Mohamed
contents Multimodal large language models (MLLMs) have demonstrated impressive cross-domain capabilities, yet their proficiency in specialized scientific fields like marine biology remains underexplored. In this work, we systematically evaluate state-of-the-art MLLMs and reveal significant limitations in their ability to perform fine-grained recognition of fish species, with the best open-source models achieving less than 10\% accuracy. This task is critical for monitoring marine ecosystems under anthropogenic pressure. To address this gap and investigate whether these failures stem from a lack of domain knowledge, we introduce FishNet++, a large-scale, multimodal benchmark. FishNet++ significantly extends existing resources with 35,133 textual descriptions for multimodal learning, 706,426 key-point annotations for morphological studies, and 119,399 bounding boxes for detection. By providing this comprehensive suite of annotations, our work facilitates the development and evaluation of specialized vision-language models capable of advancing aquatic science.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25564
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
Khan, Faizan Farooq
Radwan, Yousef
Abdelrahman, Eslam
Felemban, Abdulwahab
Mir, Aymen
Michiels, Nico K.
Temple, Andrew J.
Berumen, Michael L.
Elhoseiny, Mohamed
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) have demonstrated impressive cross-domain capabilities, yet their proficiency in specialized scientific fields like marine biology remains underexplored. In this work, we systematically evaluate state-of-the-art MLLMs and reveal significant limitations in their ability to perform fine-grained recognition of fish species, with the best open-source models achieving less than 10\% accuracy. This task is critical for monitoring marine ecosystems under anthropogenic pressure. To address this gap and investigate whether these failures stem from a lack of domain knowledge, we introduce FishNet++, a large-scale, multimodal benchmark. FishNet++ significantly extends existing resources with 35,133 textual descriptions for multimodal learning, 706,426 key-point annotations for morphological studies, and 119,399 bounding boxes for detection. By providing this comprehensive suite of annotations, our work facilitates the development and evaluation of specialized vision-language models capable of advancing aquatic science.
title FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.25564