A Visual Question Answering Method for SAR Ship: Breaking the Requirement for Multimodal Dataset Construction and Model Fine-Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Fei, Chen, Chengcheng, Chen, Hongyu, Chang, Yugang, Zeng, Weiming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MID: A Comprehensive Shore-Based Dataset for Multi-Scale Dense Ship Occlusion and Interaction Scenarios
von: Chang, Yugang, et al.
Veröffentlicht: (2024)
von: Chang, Yugang, et al.
Veröffentlicht: (2024)
Bring Remote Sensing Object Detect Into Nature Language Model: Using SFT Method
von: Wang, Fei, et al.
Veröffentlicht: (2025)
von: Wang, Fei, et al.
Veröffentlicht: (2025)
RSNet: A Light Framework for The Detection of SAR Ship Detection
von: Chen, Hongyu, et al.
Veröffentlicht: (2024)
von: Chen, Hongyu, et al.
Veröffentlicht: (2024)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
von: Zeng, Xingchen, et al.
Veröffentlicht: (2024)
von: Zeng, Xingchen, et al.
Veröffentlicht: (2024)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
von: Xue, Junxiao, et al.
Veröffentlicht: (2024)
von: Xue, Junxiao, et al.
Veröffentlicht: (2024)
Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
von: Kapuriya, Janak, et al.
Veröffentlicht: (2025)
von: Kapuriya, Janak, et al.
Veröffentlicht: (2025)
Multimodal Rationales for Explainable Visual Question Answering
von: Li, Kun, et al.
Veröffentlicht: (2024)
von: Li, Kun, et al.
Veröffentlicht: (2024)
RoadscapesQA: A Multitask, Multimodal Dataset for Visual Question Answering on Indian Roads
von: Iyer, Vijayasri, et al.
Veröffentlicht: (2026)
von: Iyer, Vijayasri, et al.
Veröffentlicht: (2026)
Fully Authentic Visual Question Answering Dataset from Online Communities
von: Chen, Chongyan, et al.
Veröffentlicht: (2023)
von: Chen, Chongyan, et al.
Veröffentlicht: (2023)
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs
von: You, Wenhao, et al.
Veröffentlicht: (2025)
von: You, Wenhao, et al.
Veröffentlicht: (2025)
Cross-modal Ship Re-Identification via Optical and SAR Imagery: A Novel Dataset and Method
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
von: Li, Yuyi, et al.
Veröffentlicht: (2025)
von: Li, Yuyi, et al.
Veröffentlicht: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
Denoising-Enhanced YOLO for Robust SAR Ship Detection
von: Zhao, Xiaojing, et al.
Veröffentlicht: (2026)
von: Zhao, Xiaojing, et al.
Veröffentlicht: (2026)
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms
von: Kabir, Raihan, et al.
Veröffentlicht: (2024)
von: Kabir, Raihan, et al.
Veröffentlicht: (2024)
Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
von: Movva, Prahitha, et al.
Veröffentlicht: (2025)
von: Movva, Prahitha, et al.
Veröffentlicht: (2025)
Improving Few-Shot Change Detection Visual Question Answering via Decision-Ambiguity-guided Reinforcement Fine-Tuning
von: Dong, Fuyu, et al.
Veröffentlicht: (2025)
von: Dong, Fuyu, et al.
Veröffentlicht: (2025)
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
von: Ma, Jie, et al.
Veröffentlicht: (2023)
von: Ma, Jie, et al.
Veröffentlicht: (2023)
NASTaR: NovaSAR Automated Ship Target Recognition Dataset
von: Hosseiny, Benyamin, et al.
Veröffentlicht: (2025)
von: Hosseiny, Benyamin, et al.
Veröffentlicht: (2025)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
von: Nguyen, Hieu Minh, et al.
Veröffentlicht: (2025)
von: Nguyen, Hieu Minh, et al.
Veröffentlicht: (2025)
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset
von: Mirzaei, Motahhare, et al.
Veröffentlicht: (2024)
von: Mirzaei, Motahhare, et al.
Veröffentlicht: (2024)
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks
von: Lee, Jusung, et al.
Veröffentlicht: (2024)
von: Lee, Jusung, et al.
Veröffentlicht: (2024)
Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
von: Zhong, Liangyu, et al.
Veröffentlicht: (2025)
von: Zhong, Liangyu, et al.
Veröffentlicht: (2025)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
von: Wang, Yuduo, et al.
Veröffentlicht: (2023)
von: Wang, Yuduo, et al.
Veröffentlicht: (2023)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
von: Zhang, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengxuan, et al.
Veröffentlicht: (2025)
Towards Fine-Grained Video Question Answering
von: Dai, Wei, et al.
Veröffentlicht: (2025)
von: Dai, Wei, et al.
Veröffentlicht: (2025)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
von: Al-Mohannadi, Aisha, et al.
Veröffentlicht: (2026)
von: Al-Mohannadi, Aisha, et al.
Veröffentlicht: (2026)
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2024)
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2024)
Lightweight SAR Ship Detection via Contrastive Distillation
von: Devasundaram, Surendar, et al.
Veröffentlicht: (2026)
von: Devasundaram, Surendar, et al.
Veröffentlicht: (2026)
Adapting Lightweight Vision Language Models for Radiological Visual Question Answering
von: Shourya, Aditya, et al.
Veröffentlicht: (2025)
von: Shourya, Aditya, et al.
Veröffentlicht: (2025)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
DriveLM: Driving with Graph Visual Question Answering
von: Sima, Chonghao, et al.
Veröffentlicht: (2023)
von: Sima, Chonghao, et al.
Veröffentlicht: (2023)
Multimodal Integration of Human-Like Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
MID: A Comprehensive Shore-Based Dataset for Multi-Scale Dense Ship Occlusion and Interaction Scenarios
von: Chang, Yugang, et al.
Veröffentlicht: (2024) -
Bring Remote Sensing Object Detect Into Nature Language Model: Using SFT Method
von: Wang, Fei, et al.
Veröffentlicht: (2025) -
RSNet: A Light Framework for The Detection of SAR Ship Detection
von: Chen, Hongyu, et al.
Veröffentlicht: (2024) -
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
von: Lee, Dosung, et al.
Veröffentlicht: (2025) -
Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
von: Zeng, Xingchen, et al.
Veröffentlicht: (2024)