MarineEval: Assessing the Marine Intelligence of Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wong, YuK-Kwan, To, Tuan-An, Zhang, Jipeng, Zheng, Ziqiang, Yeung, Sai-Kit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Boundary of GPT-4V on Marine Analysis: A Preliminary Case Study
by: Zheng, Ziqiang, et al.
Published: (2024)
by: Zheng, Ziqiang, et al.
Published: (2024)
ORCA: Object Recognition and Comprehension for Archiving Marine Species
by: Wong, Yuk-Kwan, et al.
Published: (2025)
by: Wong, Yuk-Kwan, et al.
Published: (2025)
CoralSCOP-LAT: Labeling and Analyzing Tool for Coral Reef Images with Dense Mask
by: Wong, Yuk-Kwan, et al.
Published: (2024)
by: Wong, Yuk-Kwan, et al.
Published: (2024)
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
by: Truong, Quang-Trung, et al.
Published: (2025)
by: Truong, Quang-Trung, et al.
Published: (2025)
Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
by: Vu, Tuan-Anh, et al.
Published: (2025)
by: Vu, Tuan-Anh, et al.
Published: (2025)
R4-CGQA: Retrieval-based Vision Language Models for Computer Graphics Image Quality Assessment
by: Li, Zhuangzi, et al.
Published: (2026)
by: Li, Zhuangzi, et al.
Published: (2026)
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
by: Lin, Ling, et al.
Published: (2026)
by: Lin, Ling, et al.
Published: (2026)
Mining Platoon Patterns from Traffic Videos
by: Bei, Yijun, et al.
Published: (2024)
by: Bei, Yijun, et al.
Published: (2024)
OS-W2S: An Automatic Labeling Engine for Language-Guided Open-Set Aerial Object Detection
by: Wei, Guoting, et al.
Published: (2025)
by: Wei, Guoting, et al.
Published: (2025)
Vision meets algae: A novel way for microalgae recognization and health monitor
by: Zhou, Shizheng, et al.
Published: (2022)
by: Zhou, Shizheng, et al.
Published: (2022)
YUV20K: A Complexity-Driven Benchmark and Trajectory-Aware Alignment Model for Video Camouflaged Object Detection
by: Liu, Yiyu, et al.
Published: (2026)
by: Liu, Yiyu, et al.
Published: (2026)
StyleCity: Large-Scale 3D Urban Scenes Stylization
by: Chen, Yingshu, et al.
Published: (2024)
by: Chen, Yingshu, et al.
Published: (2024)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
by: Liu, Qinghua, et al.
Published: (2025)
by: Liu, Qinghua, et al.
Published: (2025)
DarkDriving: A Real-World Day and Night Aligned Dataset for Autonomous Driving in the Dark Environment
by: Wang, Wuqi, et al.
Published: (2026)
by: Wang, Wuqi, et al.
Published: (2026)
AUTV: Creating Underwater Video Datasets with Pixel-wise Annotations
by: Truong, Quang Trung, et al.
Published: (2025)
by: Truong, Quang Trung, et al.
Published: (2025)
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
by: Chen, Xiuyuan, et al.
Published: (2023)
by: Chen, Xiuyuan, et al.
Published: (2023)
MoAI: Mixture of All Intelligence for Large Language and Vision Models
by: Lee, Byung-Kwan, et al.
Published: (2024)
by: Lee, Byung-Kwan, et al.
Published: (2024)
Mapping Urban Villages in China: Progress and Challenges
by: Cao, Rui, et al.
Published: (2025)
by: Cao, Rui, et al.
Published: (2025)
360DVO: Deep Visual Odometry for Monocular 360-Degree Camera
by: Guo, Xiaopeng, et al.
Published: (2026)
by: Guo, Xiaopeng, et al.
Published: (2026)
360VOTS: Visual Object Tracking and Segmentation in Omnidirectional Videos
by: Xu, Yinzhe, et al.
Published: (2024)
by: Xu, Yinzhe, et al.
Published: (2024)
Photo-SLAM: Real-time Simultaneous Localization and Photorealistic Mapping for Monocular, Stereo, and RGB-D Cameras
by: Huang, Huajian, et al.
Published: (2023)
by: Huang, Huajian, et al.
Published: (2023)
OmniGS: Fast Radiance Field Reconstruction using Omnidirectional Gaussian Splatting
by: Li, Longwei, et al.
Published: (2024)
by: Li, Longwei, et al.
Published: (2024)
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding
by: Nguyen-Truong, Hai, et al.
Published: (2024)
by: Nguyen-Truong, Hai, et al.
Published: (2024)
DPCD: A Quality Assessment Database for Dynamic Point Clouds
by: Liu, Yating, et al.
Published: (2025)
by: Liu, Yating, et al.
Published: (2025)
Domain-Robust Marine Plastic Detection Using Vision Models
by: Kataria, Saanvi
Published: (2025)
by: Kataria, Saanvi
Published: (2025)
3D Primitives are a Spatial Language for VLMs
by: Liu, Junze, et al.
Published: (2026)
by: Liu, Junze, et al.
Published: (2026)
Major TOM: Expandable Datasets for Earth Observation
by: Francis, Alistair, et al.
Published: (2024)
by: Francis, Alistair, et al.
Published: (2024)
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
by: Cui, Mingxuan, et al.
Published: (2026)
by: Cui, Mingxuan, et al.
Published: (2026)
K-FACE: A Large-Scale KIST Face Database in Consideration with Unconstrained Environments
by: Choi, Yeji, et al.
Published: (2021)
by: Choi, Yeji, et al.
Published: (2021)
FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting
by: Ji, Fengxian, et al.
Published: (2026)
by: Ji, Fengxian, et al.
Published: (2026)
Non-central panorama indoor dataset
by: Berenguel-Baeta, Bruno, et al.
Published: (2024)
by: Berenguel-Baeta, Bruno, et al.
Published: (2024)
LAND: A Longitudinal Analysis of Neuromorphic Datasets
by: Cohen, Gregory, et al.
Published: (2026)
by: Cohen, Gregory, et al.
Published: (2026)
TWIX: Automatically Reconstructing Structured Data from Templatized Documents
by: Lin, Yiming, et al.
Published: (2025)
by: Lin, Yiming, et al.
Published: (2025)
Research on Image Processing and Vectorization Storage Based on Garage Electronic Maps
by: Dou, Nan, et al.
Published: (2024)
by: Dou, Nan, et al.
Published: (2024)
Spatialyze: A Geospatial Video Analytics System with Spatial-Aware Optimizations
by: Kittivorawong, Chanwut, et al.
Published: (2023)
by: Kittivorawong, Chanwut, et al.
Published: (2023)
TRACER: Efficient Object Re-Identification in Networked Cameras through Adaptive Query Processing
by: Chunduri, Pramod, et al.
Published: (2025)
by: Chunduri, Pramod, et al.
Published: (2025)
VideoScoop: A Non-Traditional Domain-Independent Framework For Video Analysis
by: Billah, Hafsa
Published: (2025)
by: Billah, Hafsa
Published: (2025)
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
by: Wu, Junfeng, et al.
Published: (2025)
by: Wu, Junfeng, et al.
Published: (2025)
Towards a Flexible Scale-out Framework for Efficient Visual Data Query Processing
by: Verma, Rohit, et al.
Published: (2024)
by: Verma, Rohit, et al.
Published: (2024)
Similar Items
-
Exploring Boundary of GPT-4V on Marine Analysis: A Preliminary Case Study
by: Zheng, Ziqiang, et al.
Published: (2024) -
ORCA: Object Recognition and Comprehension for Archiving Marine Species
by: Wong, Yuk-Kwan, et al.
Published: (2025) -
CoralSCOP-LAT: Labeling and Analyzing Tool for Coral Reef Images with Dense Mask
by: Wong, Yuk-Kwan, et al.
Published: (2024) -
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
by: Truong, Quang-Trung, et al.
Published: (2025) -
Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
by: Vu, Tuan-Anh, et al.
Published: (2025)