OmniEarth: A Benchmark for Evaluating Vision-Language Models in Geospatial Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fu, Ronghao, Liu, Haoran, Zhang, Weijie, Lin, Zhiwen, Yang, Xiao, Zhang, Peng, Yang, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data
von: Wang, Fengxiang, et al.
Veröffentlicht: (2025)
von: Wang, Fengxiang, et al.
Veröffentlicht: (2025)
GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks
von: Liu, Ruixun, et al.
Veröffentlicht: (2025)
von: Liu, Ruixun, et al.
Veröffentlicht: (2025)
Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
von: Danish, Muhammad Sohail, et al.
Veröffentlicht: (2024)
von: Danish, Muhammad Sohail, et al.
Veröffentlicht: (2024)
SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding
von: Zhu, Jiashun, et al.
Veröffentlicht: (2026)
von: Zhu, Jiashun, et al.
Veröffentlicht: (2026)
CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic Segmentation
von: Gong, Ziyang, et al.
Veröffentlicht: (2024)
von: Gong, Ziyang, et al.
Veröffentlicht: (2024)
Evaluating and Benchmarking Foundation Models for Earth Observation and Geospatial AI
von: Dionelis, Nikolaos, et al.
Veröffentlicht: (2024)
von: Dionelis, Nikolaos, et al.
Veröffentlicht: (2024)
GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision
von: Sun, Lang, et al.
Veröffentlicht: (2026)
von: Sun, Lang, et al.
Veröffentlicht: (2026)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
CELLO: Causal Evaluation of Large Vision-Language Models
von: Chen, Meiqi, et al.
Veröffentlicht: (2024)
von: Chen, Meiqi, et al.
Veröffentlicht: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
OmniFashion: Towards Generalist Fashion Intelligence via Multi-Task Vision-Language Learning
von: Yang, Zhengwei, et al.
Veröffentlicht: (2026)
von: Yang, Zhengwei, et al.
Veröffentlicht: (2026)
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
von: Ying, Kaining, et al.
Veröffentlicht: (2024)
von: Ying, Kaining, et al.
Veröffentlicht: (2024)
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
von: Liu, Che, et al.
Veröffentlicht: (2025)
von: Liu, Che, et al.
Veröffentlicht: (2025)
Omni6DPose: A Benchmark and Model for Universal 6D Object Pose Estimation and Tracking
von: Zhang, Jiyao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiyao, et al.
Veröffentlicht: (2024)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
von: Fu, Teng, et al.
Veröffentlicht: (2025)
von: Fu, Teng, et al.
Veröffentlicht: (2025)
Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
von: Qian, Zhaofang, et al.
Veröffentlicht: (2025)
von: Qian, Zhaofang, et al.
Veröffentlicht: (2025)
Earth-Adapter: Bridge the Geospatial Domain Gaps with Mixture of Frequency Adaptation
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
von: Liu, Che, et al.
Veröffentlicht: (2026)
von: Liu, Che, et al.
Veröffentlicht: (2026)
SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
von: Li, Kaiyu, et al.
Veröffentlicht: (2025)
von: Li, Kaiyu, et al.
Veröffentlicht: (2025)
Long Context Transfer from Language to Vision
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2024)
CLIPose: Category-Level Object Pose Estimation with Pre-trained Vision-Language Knowledge
von: Lin, Xiao, et al.
Veröffentlicht: (2024)
von: Lin, Xiao, et al.
Veröffentlicht: (2024)
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
von: Zhang, Yiman, et al.
Veröffentlicht: (2025)
von: Zhang, Yiman, et al.
Veröffentlicht: (2025)
OmniBench: Towards The Future of Universal Omni-Language Models
von: Li, Yizhi, et al.
Veröffentlicht: (2024)
von: Li, Yizhi, et al.
Veröffentlicht: (2024)
Geospatial Foundational Embedder: Top-1 Winning Solution on EarthVision Embed2Scale Challenge (CVPR 2025)
von: Xu, Zirui, et al.
Veröffentlicht: (2025)
von: Xu, Zirui, et al.
Veröffentlicht: (2025)
OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
von: Yang, Morunliu, et al.
Veröffentlicht: (2026)
von: Yang, Morunliu, et al.
Veröffentlicht: (2026)
Evaluating Attribute Comprehension in Large Vision-Language Models
von: Zhang, Haiwen, et al.
Veröffentlicht: (2024)
von: Zhang, Haiwen, et al.
Veröffentlicht: (2024)
Vision-Language Models for Vision Tasks: A Survey
von: Zhang, Jingyi, et al.
Veröffentlicht: (2023)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2023)
DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark
von: Li, Haodong, et al.
Veröffentlicht: (2024)
von: Li, Haodong, et al.
Veröffentlicht: (2024)
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
von: Lin, Ling, et al.
Veröffentlicht: (2026)
von: Lin, Ling, et al.
Veröffentlicht: (2026)
Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming
von: Tong, Baoshun, et al.
Veröffentlicht: (2026)
von: Tong, Baoshun, et al.
Veröffentlicht: (2026)
MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
von: Shen, Hui, et al.
Veröffentlicht: (2026)
von: Shen, Hui, et al.
Veröffentlicht: (2026)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025) -
SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025) -
OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data
von: Wang, Fengxiang, et al.
Veröffentlicht: (2025) -
GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning
von: Yang, Xiao, et al.
Veröffentlicht: (2026) -
ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks
von: Liu, Ruixun, et al.
Veröffentlicht: (2025)