Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Shao, Run, Zhang, Zhaoyang, Tao, Chao, Zhang, Yunsheng, Peng, Chengli, Li, Haifeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
di: Li, Yueying, et al.
Pubblicazione: (2026)
di: Li, Yueying, et al.
Pubblicazione: (2026)
HSONet:A Siamese foreground association-driven hard case sample optimization network for high-resolution remote sensing image change detection
di: Tao, Chao, et al.
Pubblicazione: (2024)
di: Tao, Chao, et al.
Pubblicazione: (2024)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
di: Qu, Liao, et al.
Pubblicazione: (2024)
di: Qu, Liao, et al.
Pubblicazione: (2024)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
di: Chen, Xueyi, et al.
Pubblicazione: (2025)
di: Chen, Xueyi, et al.
Pubblicazione: (2025)
RS-HyRe-R1: A Hybrid Reward Mechanism to Overcome Perceptual Inertia for Remote Sensing Images Understanding
di: Zhou, Gaozhi, et al.
Pubblicazione: (2026)
di: Zhou, Gaozhi, et al.
Pubblicazione: (2026)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
di: Zheng, Anlin, et al.
Pubblicazione: (2025)
di: Zheng, Anlin, et al.
Pubblicazione: (2025)
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
di: Xu, Linrui, et al.
Pubblicazione: (2024)
di: Xu, Linrui, et al.
Pubblicazione: (2024)
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
di: Jin, Xinqi, et al.
Pubblicazione: (2025)
di: Jin, Xinqi, et al.
Pubblicazione: (2025)
Scaling Image Tokenizers with Grouped Spherical Quantization
di: Wang, Jiangtao, et al.
Pubblicazione: (2024)
di: Wang, Jiangtao, et al.
Pubblicazione: (2024)
V2Flow: Unifying Visual Tokenization and Large Language Model Vocabularies for Autoregressive Image Generation
di: Zhang, Guiwei, et al.
Pubblicazione: (2025)
di: Zhang, Guiwei, et al.
Pubblicazione: (2025)
When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
di: Luo, Junwei, et al.
Pubblicazione: (2025)
di: Luo, Junwei, et al.
Pubblicazione: (2025)
UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing
di: Dang, Yunkai, et al.
Pubblicazione: (2026)
di: Dang, Yunkai, et al.
Pubblicazione: (2026)
Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
di: Wang, Weixing, et al.
Pubblicazione: (2025)
di: Wang, Weixing, et al.
Pubblicazione: (2025)
How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A
di: Huang, YiJie, et al.
Pubblicazione: (2026)
di: Huang, YiJie, et al.
Pubblicazione: (2026)
TMCIR: Token Merge Benefits Composed Image Retrieval
di: Wang, Chaoyang, et al.
Pubblicazione: (2025)
di: Wang, Chaoyang, et al.
Pubblicazione: (2025)
Enhancing Scene Classification in Cloudy Image Scenarios: A Collaborative Transfer Method with Information Regulation Mechanism using Optical Cloud-Covered and SAR Remote Sensing Images
di: Wang, Yuze, et al.
Pubblicazione: (2025)
di: Wang, Yuze, et al.
Pubblicazione: (2025)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
di: Jiao, Yang, et al.
Pubblicazione: (2025)
di: Jiao, Yang, et al.
Pubblicazione: (2025)
Aquila-plus: Prompt-Driven Visual-Language Models for Pixel-Level Remote Sensing Image Understanding
di: Lu, Kaixuan
Pubblicazione: (2024)
di: Lu, Kaixuan
Pubblicazione: (2024)
Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification
di: Xu, Wenjia, et al.
Pubblicazione: (2024)
di: Xu, Wenjia, et al.
Pubblicazione: (2024)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
di: Zhang, Qizhe, et al.
Pubblicazione: (2024)
di: Zhang, Qizhe, et al.
Pubblicazione: (2024)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
di: Qi, Haozhe, et al.
Pubblicazione: (2026)
di: Qi, Haozhe, et al.
Pubblicazione: (2026)
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
di: Wang, Fengxiang, et al.
Pubblicazione: (2026)
di: Wang, Fengxiang, et al.
Pubblicazione: (2026)
GaussianToken: An Effective Image Tokenizer with 2D Gaussian Splatting
di: Dong, Jiajun, et al.
Pubblicazione: (2025)
di: Dong, Jiajun, et al.
Pubblicazione: (2025)
Open-Vocabulary Remote Sensing Image Semantic Segmentation
di: Cao, Qinglong, et al.
Pubblicazione: (2024)
di: Cao, Qinglong, et al.
Pubblicazione: (2024)
Hita: Holistic Tokenizer for Autoregressive Image Generation
di: Zheng, Anlin, et al.
Pubblicazione: (2025)
di: Zheng, Anlin, et al.
Pubblicazione: (2025)
Aquila: A Hierarchically Aligned Visual-Language Model for Enhanced Remote Sensing Image Comprehension
di: Lu, Kaixuan, et al.
Pubblicazione: (2024)
di: Lu, Kaixuan, et al.
Pubblicazione: (2024)
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
di: Wen, Yuhang, et al.
Pubblicazione: (2023)
di: Wen, Yuhang, et al.
Pubblicazione: (2023)
Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing
di: Zhang, Weiyu, et al.
Pubblicazione: (2026)
di: Zhang, Weiyu, et al.
Pubblicazione: (2026)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
di: Chen, Cong, et al.
Pubblicazione: (2025)
di: Chen, Cong, et al.
Pubblicazione: (2025)
QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression
di: Li, Zhongyang, et al.
Pubblicazione: (2026)
di: Li, Zhongyang, et al.
Pubblicazione: (2026)
SeFi-CD: A Semantic First Change Detection Paradigm That Can Detect Any Change You Want
di: Zhao, Ling, et al.
Pubblicazione: (2024)
di: Zhao, Ling, et al.
Pubblicazione: (2024)
Homogeneous Dynamics Space for Heterogeneous Humans
di: Liu, Xinpeng, et al.
Pubblicazione: (2024)
di: Liu, Xinpeng, et al.
Pubblicazione: (2024)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
di: Cho, Janghoon, et al.
Pubblicazione: (2025)
di: Cho, Janghoon, et al.
Pubblicazione: (2025)
Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models
di: Zhang, Jielu, et al.
Pubblicazione: (2023)
di: Zhang, Jielu, et al.
Pubblicazione: (2023)
ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling
di: Ju, Shaobo, et al.
Pubblicazione: (2026)
di: Ju, Shaobo, et al.
Pubblicazione: (2026)
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
di: Zhu, Xuanyu, et al.
Pubblicazione: (2026)
di: Zhu, Xuanyu, et al.
Pubblicazione: (2026)
On the Adversarial Robustness of Discrete Image Tokenizers
di: Bhagwatkar, Rishika, et al.
Pubblicazione: (2026)
di: Bhagwatkar, Rishika, et al.
Pubblicazione: (2026)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images
di: Li, Ke, et al.
Pubblicazione: (2025)
di: Li, Ke, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
di: Li, Yueying, et al.
Pubblicazione: (2026) -
HSONet:A Siamese foreground association-driven hard case sample optimization network for high-resolution remote sensing image change detection
di: Tao, Chao, et al.
Pubblicazione: (2024) -
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
di: Qu, Liao, et al.
Pubblicazione: (2024) -
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
di: Chen, Xueyi, et al.
Pubblicazione: (2025) -
RS-HyRe-R1: A Hybrid Reward Mechanism to Overcome Perceptual Inertia for Remote Sensing Images Understanding
di: Zhou, Gaozhi, et al.
Pubblicazione: (2026)