Exploring Boundary of GPT-4V on Marine Analysis: A Preliminary Case Study
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Ziqiang, Chen, Yiwei, Zhang, Jipeng, Vu, Tuan-Anh, Zeng, Huimin, Tim, Yue Him Wong, Yeung, Sai-Kit |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MarineEval: Assessing the Marine Intelligence of Vision-Language Models
by: Wong, YuK-Kwan, et al.
Published: (2025)
by: Wong, YuK-Kwan, et al.
Published: (2025)
Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
by: Vu, Tuan-Anh, et al.
Published: (2025)
by: Vu, Tuan-Anh, et al.
Published: (2025)
ORCA: Object Recognition and Comprehension for Archiving Marine Species
by: Wong, Yuk-Kwan, et al.
Published: (2025)
by: Wong, Yuk-Kwan, et al.
Published: (2025)
StyleCity: Large-Scale 3D Urban Scenes Stylization
by: Chen, Yingshu, et al.
Published: (2024)
by: Chen, Yingshu, et al.
Published: (2024)
CoralSCOP-LAT: Labeling and Analyzing Tool for Coral Reef Images with Dense Mask
by: Wong, Yuk-Kwan, et al.
Published: (2024)
by: Wong, Yuk-Kwan, et al.
Published: (2024)
Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
by: Vu, Tuan-Anh, et al.
Published: (2023)
by: Vu, Tuan-Anh, et al.
Published: (2023)
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding
by: Nguyen-Truong, Hai, et al.
Published: (2024)
by: Nguyen-Truong, Hai, et al.
Published: (2024)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023)
by: Van Long, Phuoc Pham, et al.
Published: (2023)
All-in-One Image Compression and Restoration
by: Zeng, Huimin, et al.
Published: (2025)
by: Zeng, Huimin, et al.
Published: (2025)
DePT3R: Joint Dense Point Tracking and 3D Reconstruction of Dynamic Scenes in a Single Forward Pass
by: Alumootil, Vivek, et al.
Published: (2025)
by: Alumootil, Vivek, et al.
Published: (2025)
OccAny: Generalized Unconstrained Urban 3D Occupancy
by: Cao, Anh-Quan, et al.
Published: (2026)
by: Cao, Anh-Quan, et al.
Published: (2026)
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
by: Truong, Quang-Trung, et al.
Published: (2025)
by: Truong, Quang-Trung, et al.
Published: (2025)
TransformMix: Learning Transformation and Mixing Strategies from Data
by: Cheung, Tsz-Him, et al.
Published: (2024)
by: Cheung, Tsz-Him, et al.
Published: (2024)
AUTV: Creating Underwater Video Datasets with Pixel-wise Annotations
by: Truong, Quang Trung, et al.
Published: (2025)
by: Truong, Quang Trung, et al.
Published: (2025)
ToXCL: A Unified Framework for Toxic Speech Detection and Explanation
by: Hoang, Nhat M., et al.
Published: (2024)
by: Hoang, Nhat M., et al.
Published: (2024)
360DVO: Deep Visual Odometry for Monocular 360-Degree Camera
by: Guo, Xiaopeng, et al.
Published: (2026)
by: Guo, Xiaopeng, et al.
Published: (2026)
360VOTS: Visual Object Tracking and Segmentation in Omnidirectional Videos
by: Xu, Yinzhe, et al.
Published: (2024)
by: Xu, Yinzhe, et al.
Published: (2024)
Photo-SLAM: Real-time Simultaneous Localization and Photorealistic Mapping for Monocular, Stereo, and RGB-D Cameras
by: Huang, Huajian, et al.
Published: (2023)
by: Huang, Huajian, et al.
Published: (2023)
OmniGS: Fast Radiance Field Reconstruction using Omnidirectional Gaussian Splatting
by: Li, Longwei, et al.
Published: (2024)
by: Li, Longwei, et al.
Published: (2024)
Facial Expression Recognition Using Residual Masking Network
by: Pham, Luan, et al.
Published: (2026)
by: Pham, Luan, et al.
Published: (2026)
More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
by: Vu, Duc Anh, et al.
Published: (2025)
by: Vu, Duc Anh, et al.
Published: (2025)
A4O: All Trigger for One sample
by: Vu, Duc Anh, et al.
Published: (2025)
by: Vu, Duc Anh, et al.
Published: (2025)
Harnessing GPT-4V(ision) for Insurance: A Preliminary Exploration
by: Lin, Chenwei, et al.
Published: (2024)
by: Lin, Chenwei, et al.
Published: (2024)
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
by: Lu, Hao, et al.
Published: (2024)
by: Lu, Hao, et al.
Published: (2024)
Open-Vocabulary Federated Learning with Multimodal Prototyping
by: Zeng, Huimin, et al.
Published: (2024)
by: Zeng, Huimin, et al.
Published: (2024)
Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment
by: Nguyen, Cong-Duy, et al.
Published: (2023)
by: Nguyen, Cong-Duy, et al.
Published: (2023)
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
by: Nguyen, Cong-Duy, et al.
Published: (2025)
by: Nguyen, Cong-Duy, et al.
Published: (2025)
Exploring the Potential of Large Language Models in Computational Argumentation
by: Chen, Guizhen, et al.
Published: (2023)
by: Chen, Guizhen, et al.
Published: (2023)
An Empirical Analysis of GPT-4V's Performance on Fashion Aesthetic Evaluation
by: Hirakawa, Yuki, et al.
Published: (2024)
by: Hirakawa, Yuki, et al.
Published: (2024)
Assessing the Effectiveness of GPT-4o in Climate Change Evidence Synthesis and Systematic Assessments: Preliminary Insights
by: Joe, Elphin Tom, et al.
Published: (2024)
by: Joe, Elphin Tom, et al.
Published: (2024)
Anti-I2V: Safeguarding your photos from malicious image-to-video generation
by: Vu, Duc, et al.
Published: (2026)
by: Vu, Duc, et al.
Published: (2026)
Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention
by: Truong, Quang-Trung, et al.
Published: (2024)
by: Truong, Quang-Trung, et al.
Published: (2024)
Color Alignment in Diffusion
by: Shum, Ka Chun, et al.
Published: (2025)
by: Shum, Ka Chun, et al.
Published: (2025)
EFHQ: Multi-purpose ExtremePose-Face-HQ dataset
by: Dao, Trung Tuan, et al.
Published: (2023)
by: Dao, Trung Tuan, et al.
Published: (2023)
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
by: Mao, Xin, et al.
Published: (2024)
by: Mao, Xin, et al.
Published: (2024)
DeID-GPT: Zero-shot Medical Text De-Identification by GPT-4
by: Liu, Zhengliang, et al.
Published: (2023)
by: Liu, Zhengliang, et al.
Published: (2023)
Arbitrary-Scale 3D Gaussian Super-Resolution
by: Zeng, Huimin, et al.
Published: (2025)
by: Zeng, Huimin, et al.
Published: (2025)
360Loc: A Dataset and Benchmark for Omnidirectional Visual Localization with Cross-device Queries
by: Huang, Huajian, et al.
Published: (2023)
by: Huang, Huajian, et al.
Published: (2023)
Towards Efficient Communication and Secure Federated Recommendation System via Low-rank Training
by: Nguyen, Ngoc-Hieu, et al.
Published: (2024)
by: Nguyen, Ngoc-Hieu, et al.
Published: (2024)
As Simple as Fine-tuning: LLM Alignment via Bidirectional Negative Feedback Loss
by: Mao, Xin, et al.
Published: (2024)
by: Mao, Xin, et al.
Published: (2024)
Similar Items
-
MarineEval: Assessing the Marine Intelligence of Vision-Language Models
by: Wong, YuK-Kwan, et al.
Published: (2025) -
Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
by: Vu, Tuan-Anh, et al.
Published: (2025) -
ORCA: Object Recognition and Comprehension for Archiving Marine Species
by: Wong, Yuk-Kwan, et al.
Published: (2025) -
StyleCity: Large-Scale 3D Urban Scenes Stylization
by: Chen, Yingshu, et al.
Published: (2024) -
CoralSCOP-LAT: Labeling and Analyzing Tool for Coral Reef Images with Dense Mask
by: Wong, Yuk-Kwan, et al.
Published: (2024)