Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yueying, Wang, Fengxiang, Li, Yan, Chen, Mingshuo, Zhao, Mengying, Lan, Long |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
by: Wang, Fengxiang, et al.
Published: (2026)
by: Wang, Fengxiang, et al.
Published: (2026)
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing Understanding
by: Wang, Fengxiang, et al.
Published: (2026)
by: Wang, Fengxiang, et al.
Published: (2026)
UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing
by: Dang, Yunkai, et al.
Published: (2026)
by: Dang, Yunkai, et al.
Published: (2026)
RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding
by: Zhu, Jiashun, et al.
Published: (2026)
by: Zhu, Jiashun, et al.
Published: (2026)
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
by: Shao, Run, et al.
Published: (2024)
by: Shao, Run, et al.
Published: (2024)
Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
by: Han, Yuhang, et al.
Published: (2024)
by: Han, Yuhang, et al.
Published: (2024)
A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
by: Dang, Yunkai, et al.
Published: (2025)
by: Dang, Yunkai, et al.
Published: (2025)
Training-Free Point Cloud Recognition Based on Geometric and Semantic Information Fusion
by: Chen, Yan, et al.
Published: (2024)
by: Chen, Yan, et al.
Published: (2024)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
by: Song, Wei, et al.
Published: (2025)
by: Song, Wei, et al.
Published: (2025)
Dual-Branch Remote Sensing Infrared Image Super-Resolution
by: Ge, Xining, et al.
Published: (2026)
by: Ge, Xining, et al.
Published: (2026)
InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition
by: Zheng, Yijie, et al.
Published: (2025)
by: Zheng, Yijie, et al.
Published: (2025)
GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search
by: Zhou, Yunqi, et al.
Published: (2025)
by: Zhou, Yunqi, et al.
Published: (2025)
Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens
by: Perron, Yohann, et al.
Published: (2026)
by: Perron, Yohann, et al.
Published: (2026)
SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering
by: Li, Wenli, et al.
Published: (2026)
by: Li, Wenli, et al.
Published: (2026)
ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG
by: Zhang, Zilun, et al.
Published: (2024)
by: Zhang, Zilun, et al.
Published: (2024)
Co-Training Vision Language Models for Remote Sensing Multi-task Learning
by: Li, Qingyun, et al.
Published: (2025)
by: Li, Qingyun, et al.
Published: (2025)
LMFNet: An Efficient Multimodal Fusion Approach for Semantic Segmentation in High-Resolution Remote Sensing
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
UHR-DETR: Efficient End-to-End Small Object Detection for Ultra-High-Resolution Remote Sensing Imagery
by: Li, Jingfang, et al.
Published: (2026)
by: Li, Jingfang, et al.
Published: (2026)
Geometric Knowledge-Assisted Federated Dual Knowledge Distillation Approach Towards Remote Sensing Satellite Imagery
by: Zou, Luyao, et al.
Published: (2026)
by: Zou, Luyao, et al.
Published: (2026)
DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation
by: Li, Boyi, et al.
Published: (2025)
by: Li, Boyi, et al.
Published: (2025)
Highly Compressed Tokenizer Can Generate Without Training
by: Beyer, L. Lao, et al.
Published: (2025)
by: Beyer, L. Lao, et al.
Published: (2025)
Structure-Semantic Decoupled Modulation of Global Geospatial Embeddings for High-Resolution Remote Sensing Mapping
by: Lyu, Jienan, et al.
Published: (2026)
by: Lyu, Jienan, et al.
Published: (2026)
Reconciling Semantic Controllability and Diversity for Remote Sensing Image Synthesis with Hybrid Semantic Embedding
by: Liu, Junde, et al.
Published: (2024)
by: Liu, Junde, et al.
Published: (2024)
GeoSeg: Training-Free Reasoning-Driven Segmentation in Remote Sensing Imagery
by: Jiang, Lifan, et al.
Published: (2026)
by: Jiang, Lifan, et al.
Published: (2026)
Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models
by: Zhang, Jielu, et al.
Published: (2023)
by: Zhang, Jielu, et al.
Published: (2023)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation
by: Yi, Chen, et al.
Published: (2025)
by: Yi, Chen, et al.
Published: (2025)
Multilingual Training-Free Remote Sensing Image Captioning
by: Rebelo, Carlos, et al.
Published: (2025)
by: Rebelo, Carlos, et al.
Published: (2025)
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
by: Yu, Hanxun, et al.
Published: (2026)
by: Yu, Hanxun, et al.
Published: (2026)
F2Net: A Frequency-Fused Network for Ultra-High Resolution Remote Sensing Segmentation
by: Chen, Hengzhi, et al.
Published: (2025)
by: Chen, Hengzhi, et al.
Published: (2025)
Semantically Robust Unsupervised Image Translation for Paired Remote Sensing Images
by: Fang, Sheng, et al.
Published: (2025)
by: Fang, Sheng, et al.
Published: (2025)
Ultra-Low Complexity On-Orbit Compression for Remote Sensing Imagery via Block Modulated Imaging
by: Wang, Zhibin, et al.
Published: (2024)
by: Wang, Zhibin, et al.
Published: (2024)
UNetMamba: An Efficient UNet-Like Mamba for Semantic Segmentation of High-Resolution Remote Sensing Images
by: Zhu, Enze, et al.
Published: (2024)
by: Zhu, Enze, et al.
Published: (2024)
Spatial-Regularization-Aware Dual-Branch Collaborative Inference for Training-Free OVSS in Remote Sensing Imagery
by: Wang, Jianzheng, et al.
Published: (2026)
by: Wang, Jianzheng, et al.
Published: (2026)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
LFTR: Learning-Free Token Reduction for Multimodal Large Language Models
by: Zhao, Zihui, et al.
Published: (2025)
by: Zhao, Zihui, et al.
Published: (2025)
Similar Items
-
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
by: Wang, Fengxiang, et al.
Published: (2026) -
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
by: Wang, Fengxiang, et al.
Published: (2025) -
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
by: Wang, Fengxiang, et al.
Published: (2025) -
Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing Understanding
by: Wang, Fengxiang, et al.
Published: (2026) -
UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing
by: Dang, Yunkai, et al.
Published: (2026)