SeafloorAI: A Large-scale Vision-Language Dataset for Seafloor Geological Survey
Fuente:
arXiv
Salvato in:
| Autori principali: | Nguyen, Kien X., Qiao, Fengchun, Trembanis, Arthur, Peng, Xi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adaptive Cascading Network for Continual Test-Time Adaptation
di: Nguyen, Kien X., et al.
Pubblicazione: (2024)
di: Nguyen, Kien X., et al.
Pubblicazione: (2024)
Single View Seafloor Recovery from Imaging Sonar via Differentiable Rendering
di: Brodjian, Sevan, et al.
Pubblicazione: (2026)
di: Brodjian, Sevan, et al.
Pubblicazione: (2026)
A Survey on Hallucination in Large Vision-Language Models
di: Liu, Hanchao, et al.
Pubblicazione: (2024)
di: Liu, Hanchao, et al.
Pubblicazione: (2024)
Large-scale Dataset Pruning with Dynamic Uncertainty
di: He, Muyang, et al.
Pubblicazione: (2023)
di: He, Muyang, et al.
Pubblicazione: (2023)
Interpretable Failure Detection with Human-Level Concepts
di: Nguyen, Kien X., et al.
Pubblicazione: (2025)
di: Nguyen, Kien X., et al.
Pubblicazione: (2025)
Vision Language Models are Biased
di: Vo, An, et al.
Pubblicazione: (2025)
di: Vo, An, et al.
Pubblicazione: (2025)
Real-time Seafloor Segmentation and Mapping
di: Grimaldi, Michele, et al.
Pubblicazione: (2025)
di: Grimaldi, Michele, et al.
Pubblicazione: (2025)
Improving Diversity in Black-box Few-shot Knowledge Distillation
di: Vo, Tri-Nhan, et al.
Pubblicazione: (2026)
di: Vo, Tri-Nhan, et al.
Pubblicazione: (2026)
Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings
di: Peng, Yunxiang, et al.
Pubblicazione: (2026)
di: Peng, Yunxiang, et al.
Pubblicazione: (2026)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
di: Li, Ang, et al.
Pubblicazione: (2025)
di: Li, Ang, et al.
Pubblicazione: (2025)
Survey of Quantization Techniques for On-Device Vision-based Crack Detection
di: Zhang, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhang, Yuxuan, et al.
Pubblicazione: (2025)
Rethinking Large-scale Dataset Compression: Shifting Focus From Labels to Images
di: Xiao, Lingao, et al.
Pubblicazione: (2025)
di: Xiao, Lingao, et al.
Pubblicazione: (2025)
ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
di: Chitty-Venkata, Krishna Teja, et al.
Pubblicazione: (2025)
di: Chitty-Venkata, Krishna Teja, et al.
Pubblicazione: (2025)
Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies
di: Pathak, Surendra, et al.
Pubblicazione: (2026)
di: Pathak, Surendra, et al.
Pubblicazione: (2026)
Diverse Image Priors for Black-box Data-free Knowledge Distillation
di: Vo, Tri-Nhan, et al.
Pubblicazione: (2026)
di: Vo, Tri-Nhan, et al.
Pubblicazione: (2026)
Graph Neural Networks in Vision-Language Image Understanding: A Survey
di: Senior, Henry, et al.
Pubblicazione: (2023)
di: Senior, Henry, et al.
Pubblicazione: (2023)
Matryoshka Query Transformer for Large Vision-Language Models
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
di: Qi, Yayun, et al.
Pubblicazione: (2024)
di: Qi, Yayun, et al.
Pubblicazione: (2024)
A Survey on Deep Clustering: From the Prior Perspective
di: Lu, Yiding, et al.
Pubblicazione: (2024)
di: Lu, Yiding, et al.
Pubblicazione: (2024)
Attention Transfer Is Not Universally Effective for Vision Transformers
di: Qin, Huaiyuan, et al.
Pubblicazione: (2026)
di: Qin, Huaiyuan, et al.
Pubblicazione: (2026)
Improved Alignment of Modalities in Large Vision Language Models
di: Jangra, Kartik, et al.
Pubblicazione: (2025)
di: Jangra, Kartik, et al.
Pubblicazione: (2025)
Detecting and Preventing Hallucinations in Large Vision Language Models
di: Gunjal, Anisha, et al.
Pubblicazione: (2023)
di: Gunjal, Anisha, et al.
Pubblicazione: (2023)
Exploiting Alpha Transparency In Language And Vision-Based AI Systems
di: Noever, David, et al.
Pubblicazione: (2024)
di: Noever, David, et al.
Pubblicazione: (2024)
StableSemantics: A Synthetic Language-Vision Dataset of Semantic Representations in Naturalistic Images
di: Zawar, Rushikesh, et al.
Pubblicazione: (2024)
di: Zawar, Rushikesh, et al.
Pubblicazione: (2024)
Visual Prompting in Multimodal Large Language Models: A Survey
di: Wu, Junda, et al.
Pubblicazione: (2024)
di: Wu, Junda, et al.
Pubblicazione: (2024)
Meply: A Large-scale Dataset and Baseline Evaluations for Metastatic Perirectal Lymph Node Detection and Segmentation
di: Guo, Weidong, et al.
Pubblicazione: (2024)
di: Guo, Weidong, et al.
Pubblicazione: (2024)
Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
Multilingual Diversity Improves Vision-Language Representations
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
Synthetic Geology: Structural Geology Meets Deep Learning
di: Ghyselincks, Simon, et al.
Pubblicazione: (2025)
di: Ghyselincks, Simon, et al.
Pubblicazione: (2025)
Yo'LLaVA: Your Personalized Language and Vision Assistant
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
CO-EVO: Co-evolving Semantic Anchoring and Style Diversification for Federated DG-ReID
di: Zhang, Fengchun, et al.
Pubblicazione: (2026)
di: Zhang, Fengchun, et al.
Pubblicazione: (2026)
Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models
di: Yang, Xu, et al.
Pubblicazione: (2023)
di: Yang, Xu, et al.
Pubblicazione: (2023)
Efficient Adaptation of Large Vision Transformer via Adapter Re-Composing
di: Dong, Wei, et al.
Pubblicazione: (2023)
di: Dong, Wei, et al.
Pubblicazione: (2023)
JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
di: Jin, Haibo, et al.
Pubblicazione: (2024)
di: Jin, Haibo, et al.
Pubblicazione: (2024)
Towards Large-scale Chemical Reaction Image Parsing via a Multimodal Large Language Model
di: Chen, Yufan, et al.
Pubblicazione: (2025)
di: Chen, Yufan, et al.
Pubblicazione: (2025)
MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification
di: Nguyen, Anh-Tien, et al.
Pubblicazione: (2025)
di: Nguyen, Anh-Tien, et al.
Pubblicazione: (2025)
Semihierarchical Reconstruction and Weak-area Revisiting for Robotic Visual Seafloor Mapping
di: She, Mengkun, et al.
Pubblicazione: (2023)
di: She, Mengkun, et al.
Pubblicazione: (2023)
Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
di: Taraday, Mitchell Keren, et al.
Pubblicazione: (2025)
di: Taraday, Mitchell Keren, et al.
Pubblicazione: (2025)
Towards Understanding How Knowledge Evolves in Large Vision-Language Models
di: Wang, Sudong, et al.
Pubblicazione: (2025)
di: Wang, Sudong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Adaptive Cascading Network for Continual Test-Time Adaptation
di: Nguyen, Kien X., et al.
Pubblicazione: (2024) -
Single View Seafloor Recovery from Imaging Sonar via Differentiable Rendering
di: Brodjian, Sevan, et al.
Pubblicazione: (2026) -
A Survey on Hallucination in Large Vision-Language Models
di: Liu, Hanchao, et al.
Pubblicazione: (2024) -
Large-scale Dataset Pruning with Dynamic Uncertainty
di: He, Muyang, et al.
Pubblicazione: (2023) -
Interpretable Failure Detection with Human-Level Concepts
di: Nguyen, Kien X., et al.
Pubblicazione: (2025)