FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
Fuente:
arXiv
Saved in:
| Main Authors: | Khan, Faizan Farooq, Radwan, Yousef, Abdelrahman, Eslam, Felemban, Abdulwahab, Mir, Aymen, Michiels, Nico K., Temple, Andrew J., Berumen, Michael L., Elhoseiny, Mohamed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
iMotion-LLM: Instruction-Conditioned Trajectory Generation
by: Felemban, Abdulwahab, et al.
Published: (2024)
by: Felemban, Abdulwahab, et al.
Published: (2024)
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024)
by: Gou, Chenhui, et al.
Published: (2024)
ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
by: Felemban, Abdulwahab, et al.
Published: (2025)
by: Felemban, Abdulwahab, et al.
Published: (2025)
FishNet: Deep Neural Networks for Low-Cost Fish Stock Estimation
by: Mots'oehli, Moseli, et al.
Published: (2024)
by: Mots'oehli, Moseli, et al.
Published: (2024)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
Neural Catalog: Scaling Species Recognition with Catalog of Life-Augmented Generation
by: Khan, Faizan Farooq, et al.
Published: (2025)
by: Khan, Faizan Farooq, et al.
Published: (2025)
Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
by: Khan, Faizan Farooq, et al.
Published: (2025)
by: Khan, Faizan Farooq, et al.
Published: (2025)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
A Shared Valence Axis Across Modern LLMs and Human EEG: The Saturation Regularity
by: Radwan, Yousef A., et al.
Published: (2026)
by: Radwan, Yousef A., et al.
Published: (2026)
Step-by-step Layered Design Generation
by: Khan, Faizan Farooq, et al.
Published: (2025)
by: Khan, Faizan Farooq, et al.
Published: (2025)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
by: Ahmed, Mahmoud, et al.
Published: (2024)
by: Ahmed, Mahmoud, et al.
Published: (2024)
ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
AI Art Neural Constellation: Revealing the Collective and Contrastive State of AI-Generated and Human Art
by: Khan, Faizan Farooq, et al.
Published: (2024)
by: Khan, Faizan Farooq, et al.
Published: (2024)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
The MATAN Model: A Reduced-Order Mathematical Framework for Volcanic Differentiation at the Matan Centre, Northern Harrat Rahat
by: Felemban, Lamees
Published: (2026)
by: Felemban, Lamees
Published: (2026)
INSTA-YOLO: Real-Time Instance Segmentation
by: Mohamed, Eslam, et al.
Published: (2021)
by: Mohamed, Eslam, et al.
Published: (2021)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
by: Shen, Xiaoqian, et al.
Published: (2023)
by: Shen, Xiaoqian, et al.
Published: (2023)
10 | THE ROLE OF ARTIFICIAL INTELLIGENCE IN CLINICAL TRIAL DESIGN AND ANALYSIS
by: S. Michiels
Published: (2025)
by: S. Michiels
Published: (2025)
CompoSE: Compositional Synthesis and Editing of 3D Shapes via Part-Aware Control
by: Slim, Habib, et al.
Published: (2026)
by: Slim, Habib, et al.
Published: (2026)
Detection of morphometric differentiation in Sattar snowtrout, Schizothorax curvifrons (Cypriniformes: Cyprinidae) from Kashmir Himalaya using a truss network system
by: Farooq Ahmad Mir
Published: (2014)
by: Farooq Ahmad Mir
Published: (2014)
Factors Driving Background Choice in Scorpionfish
by: Leonie John, et al.
Published: (2025)
by: Leonie John, et al.
Published: (2025)
Factors Driving Background Choice in Scorpionfish.
by: John, Leonie, et al.
Published: (2025)
by: John, Leonie, et al.
Published: (2025)
MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models
by: Ahmed, Seif, et al.
Published: (2025)
by: Ahmed, Seif, et al.
Published: (2025)
Generalized Replica Manifolds I: Surgery and Averaging
by: Radwan, Mohamed Hany
Published: (2025)
by: Radwan, Mohamed Hany
Published: (2025)
Simultaneous Dimensionality Reduction for Extracting Useful Representations of Large Empirical Multimodal Datasets
by: Abdelaleem, Eslam
Published: (2024)
by: Abdelaleem, Eslam
Published: (2024)
A Comparison of Recent Algorithms for Symbolic Regression to Genetic Programming
by: Radwan, Yousef A., et al.
Published: (2024)
by: Radwan, Yousef A., et al.
Published: (2024)
M-MiniGPT4: Multilingual VLLM Alignment via Translated Data
by: Han, Seung Hun, et al.
Published: (2026)
by: Han, Seung Hun, et al.
Published: (2026)
A simulation-based approach to the fluid-structure interaction inside fatigue cracks in hydraulic components
by: Michiels, Lukas
Published: (2026)
by: Michiels, Lukas
Published: (2026)
A simulation-based approach to the fluid-structure interaction inside fatigue cracks in hydraulic components
by: Michiels, Lukas
Published: (2026)
by: Michiels, Lukas
Published: (2026)
GENDER EFFECT ON CLOUD COMPUTING SERVICES ADOPTION BY UNIVERSITY STUDENTS: CASE STUDY OF SAUDI ARABIA
by: Abdulwahab Ali Almazroi
Published: (2019)
by: Abdulwahab Ali Almazroi
Published: (2019)
Commuting Toeplitz operators with biharmonic symbols
by: Bouhali, Aissa, et al.
Published: (2026)
by: Bouhali, Aissa, et al.
Published: (2026)
Commuting Toeplitz Operators With Mixed Quasihomogeneous and Analytic Symbols
by: Aissa Bouhali, et al.
Published: (2026)
by: Aissa Bouhali, et al.
Published: (2026)
A Dynamic Programming Framework for Discovering Count and Values of Multilevel Image Thresholding
by: Hegazy, Eslam, et al.
Published: (2026)
by: Hegazy, Eslam, et al.
Published: (2026)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture
by: Gamil, Mohamed, et al.
Published: (2025)
by: Gamil, Mohamed, et al.
Published: (2025)
FedPartWhole: Federated domain generalization via consistent part-whole hierarchies
by: Radwan, Ahmed, et al.
Published: (2024)
by: Radwan, Ahmed, et al.
Published: (2024)
FGR-Net:Interpretable fundus imagegradeability classification based on deepreconstruction learning
by: Khalid, Saif, et al.
Published: (2024)
by: Khalid, Saif, et al.
Published: (2024)
XProvence: Zero-Cost Multilingual Context Pruning for Retrieval-Augmented Generation
by: Mohamed, Youssef, et al.
Published: (2026)
by: Mohamed, Youssef, et al.
Published: (2026)
Migration to Microservices: A Comparative Study of Decomposition Strategies and Analysis Metrics
by: chaieb, Meryam, et al.
Published: (2024)
by: chaieb, Meryam, et al.
Published: (2024)
Similar Items
-
iMotion-LLM: Instruction-Conditioned Trajectory Generation
by: Felemban, Abdulwahab, et al.
Published: (2024) -
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024) -
ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
by: Felemban, Abdulwahab, et al.
Published: (2025) -
FishNet: Deep Neural Networks for Low-Cost Fish Stock Estimation
by: Mots'oehli, Moseli, et al.
Published: (2024) -
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023)