GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Xuran, Xiong, Zhitong, Hong, Zhongcheng, Ban, Yifang, Zhu, Xiaoxiang, Zhao, Wufan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
di: Boudras, Thomas, et al.
Pubblicazione: (2025)
di: Boudras, Thomas, et al.
Pubblicazione: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025)
di: Dai, Song, et al.
Pubblicazione: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
Clustering Guided Domain-Specific Pretrained Foundation Model Very High-Resolution Arctic Remote Sensing
di: Perera, Amal S., et al.
Pubblicazione: (2026)
di: Perera, Amal S., et al.
Pubblicazione: (2026)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
di: Han, Yudong, et al.
Pubblicazione: (2026)
di: Han, Yudong, et al.
Pubblicazione: (2026)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
di: Li, Huibin, et al.
Pubblicazione: (2025)
di: Li, Huibin, et al.
Pubblicazione: (2025)
SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
di: Liu, Sicheng, et al.
Pubblicazione: (2024)
di: Liu, Sicheng, et al.
Pubblicazione: (2024)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
di: Zeng, Zhitao, et al.
Pubblicazione: (2026)
di: Zeng, Zhitao, et al.
Pubblicazione: (2026)
Embedding-Only Uplink for Onboard Retrieval Under Shift in Remote Sensing
di: Sim, Sangcheol
Pubblicazione: (2026)
di: Sim, Sangcheol
Pubblicazione: (2026)
Towards a Generalizable Fusion Architecture for Multimodal Object Detection
di: Berjawi, Jad, et al.
Pubblicazione: (2025)
di: Berjawi, Jad, et al.
Pubblicazione: (2025)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
di: Deng, Pei, et al.
Pubblicazione: (2025)
di: Deng, Pei, et al.
Pubblicazione: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
di: Rahmatullaev, Temurbek, et al.
Pubblicazione: (2025)
di: Rahmatullaev, Temurbek, et al.
Pubblicazione: (2025)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
Maximum Temperature Prediction Using Remote Sensing Data Via Convolutional Neural Network
di: Innocenti, Lorenzo, et al.
Pubblicazione: (2024)
di: Innocenti, Lorenzo, et al.
Pubblicazione: (2024)
PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance
di: Satish, Siddarth Nilol Kundur, et al.
Pubblicazione: (2026)
di: Satish, Siddarth Nilol Kundur, et al.
Pubblicazione: (2026)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
di: Zhang, Sinin, et al.
Pubblicazione: (2026)
di: Zhang, Sinin, et al.
Pubblicazione: (2026)
Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
di: Wang, Wentao, et al.
Pubblicazione: (2025)
di: Wang, Wentao, et al.
Pubblicazione: (2025)
A Nerf-Based Color Consistency Method for Remote Sensing Images
di: Zuo, Zongcheng, et al.
Pubblicazione: (2024)
di: Zuo, Zongcheng, et al.
Pubblicazione: (2024)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
di: Gasparino, Mateus Valverde, et al.
Pubblicazione: (2024)
di: Gasparino, Mateus Valverde, et al.
Pubblicazione: (2024)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
Neighborhood Feature Pooling for Remote Sensing Image Classification
di: Nia, Fahimeh Orvati, et al.
Pubblicazione: (2025)
di: Nia, Fahimeh Orvati, et al.
Pubblicazione: (2025)
Topology-Aware Latent Diffusion for 3D Shape Generation
di: Hu, Jiangbei, et al.
Pubblicazione: (2024)
di: Hu, Jiangbei, et al.
Pubblicazione: (2024)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
di: Hu, Pan
Pubblicazione: (2025)
di: Hu, Pan
Pubblicazione: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
di: Adžemović, Momir
Pubblicazione: (2025)
di: Adžemović, Momir
Pubblicazione: (2025)
Rethinking Multimodal Point Cloud Completion: A Completion-by-Correction Perspective
di: Luo, Wang, et al.
Pubblicazione: (2025)
di: Luo, Wang, et al.
Pubblicazione: (2025)
RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
di: Ge, Junyao, et al.
Pubblicazione: (2024)
di: Ge, Junyao, et al.
Pubblicazione: (2024)
M3CAD: Towards Generic Cooperative Autonomous Driving Benchmark
di: Zhu, Morui, et al.
Pubblicazione: (2025)
di: Zhu, Morui, et al.
Pubblicazione: (2025)
SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing
di: Meng, Zi, et al.
Pubblicazione: (2026)
di: Meng, Zi, et al.
Pubblicazione: (2026)
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
di: Wang, Chaoyi, et al.
Pubblicazione: (2025)
di: Wang, Chaoyi, et al.
Pubblicazione: (2025)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
di: He, Jianxiang, et al.
Pubblicazione: (2025)
di: He, Jianxiang, et al.
Pubblicazione: (2025)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
di: Komurcu, Kursat, et al.
Pubblicazione: (2026)
di: Komurcu, Kursat, et al.
Pubblicazione: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
di: Zhu, Chenglin, et al.
Pubblicazione: (2025)
di: Zhu, Chenglin, et al.
Pubblicazione: (2025)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
di: Yang, Shan
Pubblicazione: (2026)
di: Yang, Shan
Pubblicazione: (2026)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
Towards Cognitive Collaborative Robots: Semantic-Level Integration and Explainable Control for Human-Centric Cooperation
di: Oh, Jaehong
Pubblicazione: (2025)
di: Oh, Jaehong
Pubblicazione: (2025)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
di: Asanuma, Haruka, et al.
Pubblicazione: (2025)
di: Asanuma, Haruka, et al.
Pubblicazione: (2025)
Motion Perceiver: Real-Time Occupancy Forecasting for Embedded Systems
di: Ferenczi, Bryce, et al.
Pubblicazione: (2023)
di: Ferenczi, Bryce, et al.
Pubblicazione: (2023)
Documenti analoghi
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025) -
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
di: Boudras, Thomas, et al.
Pubblicazione: (2025) -
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025) -
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
di: Tong, Jingqi, et al.
Pubblicazione: (2025) -
Clustering Guided Domain-Specific Pretrained Foundation Model Very High-Resolution Arctic Remote Sensing
di: Perera, Amal S., et al.
Pubblicazione: (2026)