FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915523970727936 |
|---|---|
| author | Khan, Faizan Farooq Radwan, Yousef Abdelrahman, Eslam Felemban, Abdulwahab Mir, Aymen Michiels, Nico K. Temple, Andrew J. Berumen, Michael L. Elhoseiny, Mohamed |
| author_facet | Khan, Faizan Farooq Radwan, Yousef Abdelrahman, Eslam Felemban, Abdulwahab Mir, Aymen Michiels, Nico K. Temple, Andrew J. Berumen, Michael L. Elhoseiny, Mohamed |
| contents | Multimodal large language models (MLLMs) have demonstrated impressive cross-domain capabilities, yet their proficiency in specialized scientific fields like marine biology remains underexplored. In this work, we systematically evaluate state-of-the-art MLLMs and reveal significant limitations in their ability to perform fine-grained recognition of fish species, with the best open-source models achieving less than 10\% accuracy. This task is critical for monitoring marine ecosystems under anthropogenic pressure. To address this gap and investigate whether these failures stem from a lack of domain knowledge, we introduce FishNet++, a large-scale, multimodal benchmark. FishNet++ significantly extends existing resources with 35,133 textual descriptions for multimodal learning, 706,426 key-point annotations for morphological studies, and 119,399 bounding boxes for detection. By providing this comprehensive suite of annotations, our work facilitates the development and evaluation of specialized vision-language models capable of advancing aquatic science. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_25564 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology Khan, Faizan Farooq Radwan, Yousef Abdelrahman, Eslam Felemban, Abdulwahab Mir, Aymen Michiels, Nico K. Temple, Andrew J. Berumen, Michael L. Elhoseiny, Mohamed Computer Vision and Pattern Recognition Multimodal large language models (MLLMs) have demonstrated impressive cross-domain capabilities, yet their proficiency in specialized scientific fields like marine biology remains underexplored. In this work, we systematically evaluate state-of-the-art MLLMs and reveal significant limitations in their ability to perform fine-grained recognition of fish species, with the best open-source models achieving less than 10\% accuracy. This task is critical for monitoring marine ecosystems under anthropogenic pressure. To address this gap and investigate whether these failures stem from a lack of domain knowledge, we introduce FishNet++, a large-scale, multimodal benchmark. FishNet++ significantly extends existing resources with 35,133 textual descriptions for multimodal learning, 706,426 key-point annotations for morphological studies, and 119,399 bounding boxes for detection. By providing this comprehensive suite of annotations, our work facilitates the development and evaluation of specialized vision-language models capable of advancing aquatic science. |
| title | FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.25564 |