LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Wentao, Zhang, Jing, Wang, Di, Zhang, Qiming, Wang, Zengmao, Du, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Boosting Semi-Supervised Object Detection in Remote Sensing Images With Active Teaching
by: Zhang, Boxuan, et al.
Published: (2024)
by: Zhang, Boxuan, et al.
Published: (2024)
TPC-ViT: Token Propagation Controller for Efficient Vision Transformer
by: Zhu, Wentao
Published: (2024)
by: Zhu, Wentao
Published: (2024)
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
by: Luo, Zhiming, et al.
Published: (2026)
by: Luo, Zhiming, et al.
Published: (2026)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
by: Ma, Xiaochen, et al.
Published: (2023)
by: Ma, Xiaochen, et al.
Published: (2023)
Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
by: Si, Dongchen, et al.
Published: (2025)
by: Si, Dongchen, et al.
Published: (2025)
Efficient Visual Transformer by Learnable Token Merging
by: Wang, Yancheng, et al.
Published: (2024)
by: Wang, Yancheng, et al.
Published: (2024)
ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing
by: Ma, Xingwei, et al.
Published: (2025)
by: Ma, Xingwei, et al.
Published: (2025)
MambaHSI: Spatial-Spectral Mamba for Hyperspectral Image Classification
by: Li, Yapeng, et al.
Published: (2025)
by: Li, Yapeng, et al.
Published: (2025)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
by: Yuan, Bo, et al.
Published: (2024)
by: Yuan, Bo, et al.
Published: (2024)
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
by: Xu, Xuwei, et al.
Published: (2023)
by: Xu, Xuwei, et al.
Published: (2023)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
by: Zhu, Chen, et al.
Published: (2025)
by: Zhu, Chen, et al.
Published: (2025)
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
by: Zhang, Tianfang, et al.
Published: (2024)
by: Zhang, Tianfang, et al.
Published: (2024)
S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing
by: Lv, Liang, et al.
Published: (2025)
by: Lv, Liang, et al.
Published: (2025)
Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
by: Chen, Haoyang, et al.
Published: (2026)
by: Chen, Haoyang, et al.
Published: (2026)
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
by: Shao, Run, et al.
Published: (2024)
by: Shao, Run, et al.
Published: (2024)
ERVD: An Efficient and Robust ViT-Based Distillation Framework for Remote Sensing Image Retrieval
by: Dong, Le, et al.
Published: (2024)
by: Dong, Le, et al.
Published: (2024)
BAFNet: Bilateral Attention Fusion Network for Lightweight Semantic Segmentation of Urban Remote Sensing Images
by: Wang, Wentao, et al.
Published: (2024)
by: Wang, Wentao, et al.
Published: (2024)
GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
Boosting Multimodal Remote Sensing Image Classification with Transformer-based Heterogeneously Salient Graph Representation
by: Yang, Jiaqi, et al.
Published: (2023)
by: Yang, Jiaqi, et al.
Published: (2023)
Slender Object Scene Segmentation in Remote Sensing Image Based on Learnable Morphological Skeleton with Segment Anything Model
by: Xie, Jun, et al.
Published: (2024)
by: Xie, Jun, et al.
Published: (2024)
Exploring Text-Guided Single Image Editing for Remote Sensing Images
by: Han, Fangzhou, et al.
Published: (2024)
by: Han, Fangzhou, et al.
Published: (2024)
SAT3D: Image-driven Semantic Attribute Transfer in 3D
by: Zhai, Zhijun, et al.
Published: (2024)
by: Zhai, Zhijun, et al.
Published: (2024)
Research on Improved U-net Based Remote Sensing Image Segmentation Algorithm
by: Yang, Qiming, et al.
Published: (2024)
by: Yang, Qiming, et al.
Published: (2024)
Panoptic Perception: A Novel Task and Fine-grained Dataset for Universal Remote Sensing Image Interpretation
by: Zhao, Danpei, et al.
Published: (2024)
by: Zhao, Danpei, et al.
Published: (2024)
Pattern Integration and Enhancement Vision Transformer for Self-Supervised Learning in Remote Sensing
by: Lu, Kaixuan, et al.
Published: (2024)
by: Lu, Kaixuan, et al.
Published: (2024)
MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining
by: Wang, Di, et al.
Published: (2024)
by: Wang, Di, et al.
Published: (2024)
Multi-Tailed Vision Transformer for Efficient Inference
by: Wang, Yunke, et al.
Published: (2022)
by: Wang, Yunke, et al.
Published: (2022)
CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic Segmentation
by: Gong, Ziyang, et al.
Published: (2024)
by: Gong, Ziyang, et al.
Published: (2024)
RescueADI: Adaptive Disaster Interpretation in Remote Sensing Images with Autonomous Agents
by: Liu, Zhuoran, et al.
Published: (2024)
by: Liu, Zhuoran, et al.
Published: (2024)
TEFormer: Texture-Aware and Edge-Guided Transformer for Semantic Segmentation of Urban Remote Sensing Images
by: Zhou, Guoyu, et al.
Published: (2025)
by: Zhou, Guoyu, et al.
Published: (2025)
ViT-5: Vision Transformers for The Mid-2020s
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
Efficient Prompt Tuning of Large Vision-Language Model for Fine-Grained Ship Classification
by: Lan, Long, et al.
Published: (2024)
by: Lan, Long, et al.
Published: (2024)
Multimodal Interpretation of Remote Sensing Images: Dynamic Resolution Input Strategy and Multi-scale Vision-Language Alignment Mechanism
by: Zhang, Siyu, et al.
Published: (2025)
by: Zhang, Siyu, et al.
Published: (2025)
Reconciling Semantic Controllability and Diversity for Remote Sensing Image Synthesis with Hybrid Semantic Embedding
by: Liu, Junde, et al.
Published: (2024)
by: Liu, Junde, et al.
Published: (2024)
Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
by: Guo, Haonan, et al.
Published: (2024)
by: Guo, Haonan, et al.
Published: (2024)
APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers
by: Wu, Zhuguanyu, et al.
Published: (2025)
by: Wu, Zhuguanyu, et al.
Published: (2025)
MetaSegNet: Metadata-collaborative Vision-Language Representation Learning for Semantic Segmentation of Remote Sensing Images
by: Wang, Libo, et al.
Published: (2023)
by: Wang, Libo, et al.
Published: (2023)
ViTCN: Vision Transformer Contrastive Network For Reasoning
by: Song, Bo, et al.
Published: (2024)
by: Song, Bo, et al.
Published: (2024)
Similar Items
-
Boosting Semi-Supervised Object Detection in Remote Sensing Images With Active Teaching
by: Zhang, Boxuan, et al.
Published: (2024) -
TPC-ViT: Token Propagation Controller for Efficient Vision Transformer
by: Zhu, Wentao
Published: (2024) -
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
by: Luo, Zhiming, et al.
Published: (2026) -
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
by: Ma, Xiaochen, et al.
Published: (2023) -
Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation
by: Zhang, Jing, et al.
Published: (2026)