VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Xiang, Ding, Jian, Elhoseiny, Mohamed |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis
par: Pang, Chao, et autres
Publié: (2024)
par: Pang, Chao, et autres
Publié: (2024)
DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark
par: Li, Haodong, et autres
Publié: (2024)
par: Li, Haodong, et autres
Publié: (2024)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
par: Ataallah, Kirolos, et autres
Publié: (2024)
par: Ataallah, Kirolos, et autres
Publié: (2024)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
par: Ahmed, Mahmoud, et autres
Publié: (2025)
par: Ahmed, Mahmoud, et autres
Publié: (2025)
Continual Vision-Language Learning for Remote Sensing: Benchmarking and Analysis
par: Weng, Xingxing, et autres
Publié: (2026)
par: Weng, Xingxing, et autres
Publié: (2026)
GAIA: A Global, Multi-modal, Multi-scale Vision-Language Dataset for Remote Sensing Image Analysis
par: Zavras, Angelos, et autres
Publié: (2025)
par: Zavras, Angelos, et autres
Publié: (2025)
RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
par: Zi, Xing, et autres
Publié: (2025)
par: Zi, Xing, et autres
Publié: (2025)
SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding
par: Luo, Junwei, et autres
Publié: (2024)
par: Luo, Junwei, et autres
Publié: (2024)
A Benchmark for Multi-Lingual Vision-Language Learning in Remote Sensing Image Captioning
par: Zhou, Qing, et autres
Publié: (2025)
par: Zhou, Qing, et autres
Publié: (2025)
A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark
par: Yang, Zhigang, et autres
Publié: (2025)
par: Yang, Zhigang, et autres
Publié: (2025)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
par: Shen, Xiaoqian, et autres
Publié: (2023)
par: Shen, Xiaoqian, et autres
Publié: (2023)
Pansharpening for Thin-Cloud Contaminated Remote Sensing Images: A Unified Framework and Benchmark Dataset
par: Du, Songcheng, et autres
Publié: (2026)
par: Du, Songcheng, et autres
Publié: (2026)
Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
par: Weng, Xingxing, et autres
Publié: (2025)
par: Weng, Xingxing, et autres
Publié: (2025)
Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing
par: Anderson, Madeline, et autres
Publié: (2025)
par: Anderson, Madeline, et autres
Publié: (2025)
Hyperspectral Remote Sensing Images Salient Object Detection: The First Benchmark Dataset and Baseline
par: Liu, Peifu, et autres
Publié: (2025)
par: Liu, Peifu, et autres
Publié: (2025)
CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models
par: An, Xiao, et autres
Publié: (2024)
par: An, Xiao, et autres
Publié: (2024)
Few-Shot Adaptation Benchmark for Remote Sensing Vision-Language Models
par: Khoury, Karim El, et autres
Publié: (2025)
par: Khoury, Karim El, et autres
Publié: (2025)
ATRNet-STAR: A Large Dataset and Benchmark Towards Remote Sensing Object Recognition in the Wild
par: Liu, Yongxiang, et autres
Publié: (2025)
par: Liu, Yongxiang, et autres
Publié: (2025)
Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction
par: Chen, Jun, et autres
Publié: (2022)
par: Chen, Jun, et autres
Publié: (2022)
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
par: Ataallah, Kirolos, et autres
Publié: (2024)
par: Ataallah, Kirolos, et autres
Publié: (2024)
How Well Can Vision Language Models See Image Details?
par: Gou, Chenhui, et autres
Publié: (2024)
par: Gou, Chenhui, et autres
Publié: (2024)
RSDehamba: Lightweight Vision Mamba for Remote Sensing Satellite Image Dehazing
par: Zhou, Huiling, et autres
Publié: (2024)
par: Zhou, Huiling, et autres
Publié: (2024)
CBEN -- A Multimodal Machine Learning Dataset for Cloud Robust Remote Sensing Image Understanding
par: Stricker, Marco, et autres
Publié: (2026)
par: Stricker, Marco, et autres
Publié: (2026)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
par: Ataallah, Kirolos, et autres
Publié: (2024)
par: Ataallah, Kirolos, et autres
Publié: (2024)
Vision-Language Models in Remote Sensing: Current Progress and Future Trends
par: Li, Xiang, et autres
Publié: (2023)
par: Li, Xiang, et autres
Publié: (2023)
Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation
par: He, Yiguo, et autres
Publié: (2025)
par: He, Yiguo, et autres
Publié: (2025)
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing
par: Liu, Fan, et autres
Publié: (2023)
par: Liu, Fan, et autres
Publié: (2023)
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
par: Luo, Zhiming, et autres
Publié: (2026)
par: Luo, Zhiming, et autres
Publié: (2026)
Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning
par: Tosato, Lucrezia, et autres
Publié: (2026)
par: Tosato, Lucrezia, et autres
Publié: (2026)
GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding
par: Zhou, Yue, et autres
Publié: (2024)
par: Zhou, Yue, et autres
Publié: (2024)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
par: Shen, Xiaoqian, et autres
Publié: (2025)
par: Shen, Xiaoqian, et autres
Publié: (2025)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
par: Ahmed, Mahmoud, et autres
Publié: (2024)
par: Ahmed, Mahmoud, et autres
Publié: (2024)
Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
par: Khan, Faizan Farooq, et autres
Publié: (2025)
par: Khan, Faizan Farooq, et autres
Publié: (2025)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
par: Wang, Yuxuan, et autres
Publié: (2024)
par: Wang, Yuxuan, et autres
Publié: (2024)
Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
par: Li, Yuxuan, et autres
Publié: (2026)
par: Li, Yuxuan, et autres
Publié: (2026)
Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
par: Zhang, Xu, et autres
Publié: (2025)
par: Zhang, Xu, et autres
Publié: (2025)
SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
par: Si, Dongchen, et autres
Publié: (2025)
par: Si, Dongchen, et autres
Publié: (2025)
GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing
par: Ma, Xianzhi, et autres
Publié: (2025)
par: Ma, Xianzhi, et autres
Publié: (2025)
Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image Classification
par: Zhang, Junjie, et autres
Publié: (2025)
par: Zhang, Junjie, et autres
Publié: (2025)
RRSIS: Referring Remote Sensing Image Segmentation
par: Yuan, Zhenghang, et autres
Publié: (2023)
par: Yuan, Zhenghang, et autres
Publié: (2023)
Documents similaires
-
VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis
par: Pang, Chao, et autres
Publié: (2024) -
DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark
par: Li, Haodong, et autres
Publié: (2024) -
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
par: Ataallah, Kirolos, et autres
Publié: (2024) -
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
par: Ahmed, Mahmoud, et autres
Publié: (2025) -
Continual Vision-Language Learning for Remote Sensing: Benchmarking and Analysis
par: Weng, Xingxing, et autres
Publié: (2026)