XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Fengxiang, Wang, Hongzhen, Chen, Mingshuo, Wang, Di, Wang, Yulin, Guo, Zonghao, Ma, Qiang, Lan, Long, Yang, Wenjing, Zhang, Jing, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
di: Wang, Fengxiang, et al.
Pubblicazione: (2025)
di: Wang, Fengxiang, et al.
Pubblicazione: (2025)
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
di: Wang, Fengxiang, et al.
Pubblicazione: (2026)
di: Wang, Fengxiang, et al.
Pubblicazione: (2026)
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
di: Li, Yueying, et al.
Pubblicazione: (2026)
di: Li, Yueying, et al.
Pubblicazione: (2026)
RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing
di: Wang, Fengxiang, et al.
Pubblicazione: (2025)
di: Wang, Fengxiang, et al.
Pubblicazione: (2025)
Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling
di: Wang, Fengxiang, et al.
Pubblicazione: (2024)
di: Wang, Fengxiang, et al.
Pubblicazione: (2024)
Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing Understanding
di: Wang, Fengxiang, et al.
Pubblicazione: (2026)
di: Wang, Fengxiang, et al.
Pubblicazione: (2026)
Diffusion Enhancement for Cloud Removal in Ultra-Resolution Remote Sensing Imagery
di: Sui, Jialu, et al.
Pubblicazione: (2024)
di: Sui, Jialu, et al.
Pubblicazione: (2024)
GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
di: Zhang, Peirong, et al.
Pubblicazione: (2025)
di: Zhang, Peirong, et al.
Pubblicazione: (2025)
ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG
di: Zhang, Zilun, et al.
Pubblicazione: (2024)
di: Zhang, Zilun, et al.
Pubblicazione: (2024)
OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data
di: Wang, Fengxiang, et al.
Pubblicazione: (2025)
di: Wang, Fengxiang, et al.
Pubblicazione: (2025)
AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
di: Gong, Kaixiong, et al.
Pubblicazione: (2024)
di: Gong, Kaixiong, et al.
Pubblicazione: (2024)
Ultra-Low Complexity On-Orbit Compression for Remote Sensing Imagery via Block Modulated Imaging
di: Wang, Zhibin, et al.
Pubblicazione: (2024)
di: Wang, Zhibin, et al.
Pubblicazione: (2024)
Multimodal Urban Areas of Interest Generation via Remote Sensing Imagery and Geographical Prior
di: Shi, Chuanji, et al.
Pubblicazione: (2024)
di: Shi, Chuanji, et al.
Pubblicazione: (2024)
LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation
di: Wang, Jun, et al.
Pubblicazione: (2026)
di: Wang, Jun, et al.
Pubblicazione: (2026)
A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
di: Dang, Yunkai, et al.
Pubblicazione: (2025)
di: Dang, Yunkai, et al.
Pubblicazione: (2025)
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
di: Zhang, Yi-Fan, et al.
Pubblicazione: (2024)
di: Zhang, Yi-Fan, et al.
Pubblicazione: (2024)
UHR-DETR: Efficient End-to-End Small Object Detection for Ultra-High-Resolution Remote Sensing Imagery
di: Li, Jingfang, et al.
Pubblicazione: (2026)
di: Li, Jingfang, et al.
Pubblicazione: (2026)
ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery
di: Li, Ke, et al.
Pubblicazione: (2026)
di: Li, Ke, et al.
Pubblicazione: (2026)
LMFNet: An Efficient Multimodal Fusion Approach for Semantic Segmentation in High-Resolution Remote Sensing
di: Wang, Tong, et al.
Pubblicazione: (2024)
di: Wang, Tong, et al.
Pubblicazione: (2024)
MERLIN: Building Low-SNR Robust Multimodal LLMs for Electromagnetic Signals
di: Shen, Junyu, et al.
Pubblicazione: (2026)
di: Shen, Junyu, et al.
Pubblicazione: (2026)
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
di: Xu, Linrui, et al.
Pubblicazione: (2024)
di: Xu, Linrui, et al.
Pubblicazione: (2024)
Built Environment Reasoning from Remote Sensing Imagery Using Large Vision--Language Models
di: Wang, Dongdong, et al.
Pubblicazione: (2026)
di: Wang, Dongdong, et al.
Pubblicazione: (2026)
Assessing the Impacts of Extreme Weather Events on Photovoltaic Installations Using Remote Sensing Imagery
di: Kirsten Perry, et al.
Pubblicazione: (2025)
di: Kirsten Perry, et al.
Pubblicazione: (2025)
Learning a Cross-modality Anomaly Detector for Remote Sensing Imagery
di: Li, Jingtao, et al.
Pubblicazione: (2023)
di: Li, Jingtao, et al.
Pubblicazione: (2023)
Automated Detection of Hillforts in Remote Sensing Imagery With Deep Multimodal Segmentation
di: Daniel Canedo, et al.
Pubblicazione: (2024)
di: Daniel Canedo, et al.
Pubblicazione: (2024)
Attention! Your Vision Language Model Could Be Maliciously Manipulated
di: Wang, Xiaosen, et al.
Pubblicazione: (2025)
di: Wang, Xiaosen, et al.
Pubblicazione: (2025)
Efficient Prompt Tuning of Large Vision-Language Model for Fine-Grained Ship Classification
di: Lan, Long, et al.
Pubblicazione: (2024)
di: Lan, Long, et al.
Pubblicazione: (2024)
GatheringSense: AI-Generated Imagery and Embodied Experiences for Understanding Literati Gatherings
di: Zhou, You, et al.
Pubblicazione: (2026)
di: Zhou, You, et al.
Pubblicazione: (2026)
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding
di: Zhu, Jiashun, et al.
Pubblicazione: (2026)
di: Zhu, Jiashun, et al.
Pubblicazione: (2026)
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
di: Liu, Yi, et al.
Pubblicazione: (2026)
di: Liu, Yi, et al.
Pubblicazione: (2026)
SOPSeg: Prompt-based Small Object Instance Segmentation in Remote Sensing Imagery
di: Wang, Chenhao, et al.
Pubblicazione: (2025)
di: Wang, Chenhao, et al.
Pubblicazione: (2025)
L2RSI: Cross-view LiDAR-based Place Recognition for Large-scale Urban Scenes via Remote Sensing Imagery
di: Shi, Ziwei, et al.
Pubblicazione: (2025)
di: Shi, Ziwei, et al.
Pubblicazione: (2025)
Spatial-Regularization-Aware Dual-Branch Collaborative Inference for Training-Free OVSS in Remote Sensing Imagery
di: Wang, Jianzheng, et al.
Pubblicazione: (2026)
di: Wang, Jianzheng, et al.
Pubblicazione: (2026)
Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration
di: Li, Zhili, et al.
Pubblicazione: (2026)
di: Li, Zhili, et al.
Pubblicazione: (2026)
AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans
di: Xie, Wei, et al.
Pubblicazione: (2025)
di: Xie, Wei, et al.
Pubblicazione: (2025)
Map-Assisted Remote-Sensing Image Compression at Extremely Low Bitrates
di: Ye, Yixuan, et al.
Pubblicazione: (2024)
di: Ye, Yixuan, et al.
Pubblicazione: (2024)
ISWSST: Index-space-wave State Superposition Transformers for Multispectral Remotely Sensed Imagery Semantic Segmentation
di: Li, Chang, et al.
Pubblicazione: (2024)
di: Li, Chang, et al.
Pubblicazione: (2024)
Super-Resolution for Remote Sensing Imagery via the Coupling of a Variational Model and Deep Learning
di: Sun, Jing, et al.
Pubblicazione: (2024)
di: Sun, Jing, et al.
Pubblicazione: (2024)
Referring Change Detection in Remote Sensing Imagery
di: Korkmaz, Yilmaz, et al.
Pubblicazione: (2025)
di: Korkmaz, Yilmaz, et al.
Pubblicazione: (2025)
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
di: Luo, Zhiming, et al.
Pubblicazione: (2026)
di: Luo, Zhiming, et al.
Pubblicazione: (2026)
Documenti analoghi
-
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
di: Wang, Fengxiang, et al.
Pubblicazione: (2025) -
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
di: Wang, Fengxiang, et al.
Pubblicazione: (2026) -
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
di: Li, Yueying, et al.
Pubblicazione: (2026) -
RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing
di: Wang, Fengxiang, et al.
Pubblicazione: (2025) -
Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling
di: Wang, Fengxiang, et al.
Pubblicazione: (2024)