AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Alawode, Basit, Ganapathi, Iyyakutti Iyappan, Javed, Sajid, Werghi, Naoufel, Bennamoun, Mohammed, Mahmood, Arif |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
von: Albastaki, Shahad, et al.
Veröffentlicht: (2025)
von: Albastaki, Shahad, et al.
Veröffentlicht: (2025)
CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment
von: Javed, Sajid, et al.
Veröffentlicht: (2024)
von: Javed, Sajid, et al.
Veröffentlicht: (2024)
Predicting the Best of N Visual Trackers
von: Alawode, Basit, et al.
Veröffentlicht: (2024)
von: Alawode, Basit, et al.
Veröffentlicht: (2024)
CLDTracker: A Comprehensive Language Description for Visual Tracking
von: Alansari, Mohamad, et al.
Veröffentlicht: (2025)
von: Alansari, Mohamad, et al.
Veröffentlicht: (2025)
DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation
von: Assefa, Maregu, et al.
Veröffentlicht: (2025)
von: Assefa, Maregu, et al.
Veröffentlicht: (2025)
MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding
von: Alawode, Basit, et al.
Veröffentlicht: (2026)
von: Alawode, Basit, et al.
Veröffentlicht: (2026)
Advancing Histopathology with Deep Learning Under Data Scarcity: A Decade in Review
von: Obeid, Ahmad, et al.
Veröffentlicht: (2024)
von: Obeid, Ahmad, et al.
Veröffentlicht: (2024)
Rethinking Memory Design in SAM-Based Visual Object Tracking
von: Alansari, Mohamad, et al.
Veröffentlicht: (2025)
von: Alansari, Mohamad, et al.
Veröffentlicht: (2025)
Video Anomaly Detection in 10 Years: A Survey and Outlook
von: Abdalla, Moshira, et al.
Veröffentlicht: (2024)
von: Abdalla, Moshira, et al.
Veröffentlicht: (2024)
Implicit to Explicit Entropy Regularization: Benchmarking ViT Fine-tuning under Noisy Labels
von: Marrium, Maria, et al.
Veröffentlicht: (2024)
von: Marrium, Maria, et al.
Veröffentlicht: (2024)
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
von: Alansari, Mohamad, et al.
Veröffentlicht: (2026)
von: Alansari, Mohamad, et al.
Veröffentlicht: (2026)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
von: Zhang, Xian, et al.
Veröffentlicht: (2025)
von: Zhang, Xian, et al.
Veröffentlicht: (2025)
BENet: A Cross-domain Robust Network for Detecting Face Forgeries via Bias Expansion and Latent-space Attention
von: Liu, Weihua, et al.
Veröffentlicht: (2024)
von: Liu, Weihua, et al.
Veröffentlicht: (2024)
Multi-Modal Attention Networks for Enhanced Segmentation and Depth Estimation of Subsurface Defects in Pulse Thermography
von: Salah, Mohammed, et al.
Veröffentlicht: (2025)
von: Salah, Mohammed, et al.
Veröffentlicht: (2025)
Transformer-Based Wireless Capsule Endoscopy Bleeding Tissue Detection and Classification
von: Alawode, Basit, et al.
Veröffentlicht: (2024)
von: Alawode, Basit, et al.
Veröffentlicht: (2024)
STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
von: Velayudhan, Divya, et al.
Veröffentlicht: (2025)
von: Velayudhan, Divya, et al.
Veröffentlicht: (2025)
MUOT_3M: A 3 Million Frame Multimodal Underwater Benchmark and the MUTrack Tracking Method
von: Bakht, Ahsan Baidar, et al.
Veröffentlicht: (2026)
von: Bakht, Ahsan Baidar, et al.
Veröffentlicht: (2026)
Cytoplasmic Strings Analysis in Human Embryo Time-Lapse Videos using Deep Learning Framework
von: Sohail, Anabia, et al.
Veröffentlicht: (2025)
von: Sohail, Anabia, et al.
Veröffentlicht: (2025)
RobMOT: Robust 3D Multi-Object Tracking by Observational Noise and State Estimation Drift Mitigation on LiDAR PointCloud
von: Nagy, Mohamed, et al.
Veröffentlicht: (2024)
von: Nagy, Mohamed, et al.
Veröffentlicht: (2024)
Towards Accurate State Estimation: Kalman Filter Incorporating Motion Dynamics for 3D Multi-Object Tracking
von: Nagy, Mohamed, et al.
Veröffentlicht: (2025)
von: Nagy, Mohamed, et al.
Veröffentlicht: (2025)
Aquatic-GS: A Hybrid 3D Representation for Underwater Scenes
von: Liu, Shaohua, et al.
Veröffentlicht: (2024)
von: Liu, Shaohua, et al.
Veröffentlicht: (2024)
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing
von: Liu, Fan, et al.
Veröffentlicht: (2023)
von: Liu, Fan, et al.
Veröffentlicht: (2023)
DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation
von: Xiong, Zhitong, et al.
Veröffentlicht: (2025)
von: Xiong, Zhitong, et al.
Veröffentlicht: (2025)
AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
von: Imam, Raza, et al.
Veröffentlicht: (2024)
von: Imam, Raza, et al.
Veröffentlicht: (2024)
Face Pyramid Vision Transformer
von: Islam, Khawar, et al.
Veröffentlicht: (2022)
von: Islam, Khawar, et al.
Veröffentlicht: (2022)
Unveiling the Underwater World: CLIP Perception Model-Guided Underwater Image Enhancement
von: Cao, Jiangzhong, et al.
Veröffentlicht: (2025)
von: Cao, Jiangzhong, et al.
Veröffentlicht: (2025)
CropVLM: A Domain-Adapted Vision-Language Model for Open-Set Crop Analysis
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
Semantically-aware Neural Radiance Fields for Visual Scene Understanding: A Comprehensive Review
von: Nguyen, Thang-Anh-Quan, et al.
Veröffentlicht: (2024)
von: Nguyen, Thang-Anh-Quan, et al.
Veröffentlicht: (2024)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
von: Condez, Ana Carolina, et al.
Veröffentlicht: (2025)
von: Condez, Ana Carolina, et al.
Veröffentlicht: (2025)
VisionCLIP: An Med-AIGC based Ethical Language-Image Foundation Model for Generalizable Retina Image Analysis
von: Wei, Hao, et al.
Veröffentlicht: (2024)
von: Wei, Hao, et al.
Veröffentlicht: (2024)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
Underwater Image Enhancement by Diffusion Model with Customized CLIP-Classifier
von: Liu, Shuaixin, et al.
Veröffentlicht: (2024)
von: Liu, Shuaixin, et al.
Veröffentlicht: (2024)
Neuromorphic Vision-based Motion Segmentation with Graph Transformer Neural Network
von: Alkendi, Yusra, et al.
Veröffentlicht: (2024)
von: Alkendi, Yusra, et al.
Veröffentlicht: (2024)
ScenarioCLIP: Pretrained Transferable Visual Language Models and Action-Genome Dataset for Natural Scene Analysis
von: Sinha, Advik, et al.
Veröffentlicht: (2025)
von: Sinha, Advik, et al.
Veröffentlicht: (2025)
SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
Language Model Guided Interpretable Video Action Reasoning
von: Wang, Ning, et al.
Veröffentlicht: (2024)
von: Wang, Ning, et al.
Veröffentlicht: (2024)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
von: Albastaki, Shahad, et al.
Veröffentlicht: (2025) -
CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment
von: Javed, Sajid, et al.
Veröffentlicht: (2024) -
Predicting the Best of N Visual Trackers
von: Alawode, Basit, et al.
Veröffentlicht: (2024) -
CLDTracker: A Comprehensive Language Description for Visual Tracking
von: Alansari, Mohamad, et al.
Veröffentlicht: (2025) -
DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation
von: Assefa, Maregu, et al.
Veröffentlicht: (2025)