Aquila-plus: Prompt-Driven Visual-Language Models for Pixel-Level Remote Sensing Image Understanding
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Lu, Kaixuan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Aquila: A Hierarchically Aligned Visual-Language Model for Enhanced Remote Sensing Image Comprehension
par: Lu, Kaixuan, et autres
Publié: (2024)
par: Lu, Kaixuan, et autres
Publié: (2024)
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
par: Shao, Run, et autres
Publié: (2024)
par: Shao, Run, et autres
Publié: (2024)
SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models
par: Zhong, Chen, et autres
Publié: (2026)
par: Zhong, Chen, et autres
Publié: (2026)
Beyond Pixels: A Training-Free, Text-to-Text Framework for Remote Sensing Image Retrieval
par: Xiao, J., et autres
Publié: (2025)
par: Xiao, J., et autres
Publié: (2025)
PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval
par: Pan, Jiancheng, et autres
Publié: (2024)
par: Pan, Jiancheng, et autres
Publié: (2024)
From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing
par: Sun, Xintian, et autres
Publié: (2024)
par: Sun, Xintian, et autres
Publié: (2024)
PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding
par: Wang, Nan, et autres
Publié: (2026)
par: Wang, Nan, et autres
Publié: (2026)
Few-Shot Remote Sensing Image Scene Classification with CLIP and Prompt Learning
par: Dimitrovski, Ivica, et autres
Publié: (2025)
par: Dimitrovski, Ivica, et autres
Publié: (2025)
Direction-Oriented Visual-semantic Embedding Model for Remote Sensing Image-text Retrieval
par: Ma, Qing, et autres
Publié: (2023)
par: Ma, Qing, et autres
Publié: (2023)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
par: Liu, Ye, et autres
Publié: (2025)
par: Liu, Ye, et autres
Publié: (2025)
SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding
par: Luo, Junwei, et autres
Publié: (2024)
par: Luo, Junwei, et autres
Publié: (2024)
RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering
par: Lin, Hui, et autres
Publié: (2024)
par: Lin, Hui, et autres
Publié: (2024)
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications
par: Rahman, Ben
Publié: (2025)
par: Rahman, Ben
Publié: (2025)
CroBIM-U: Uncertainty-Driven Referring Remote Sensing Image Segmentation
par: Sun, Yuzhe, et autres
Publié: (2026)
par: Sun, Yuzhe, et autres
Publié: (2026)
Vision-Language Model Purified Semi-Supervised Semantic Segmentation for Remote Sensing Images
par: Wang, Shanwen, et autres
Publié: (2026)
par: Wang, Shanwen, et autres
Publié: (2026)
A High-Level Survey of Optical Remote Sensing
par: Koletsis, Panagiotis, et autres
Publié: (2026)
par: Koletsis, Panagiotis, et autres
Publié: (2026)
RSGen: Enhancing Layout-Driven Remote Sensing Image Generation with Diverse Edge Guidance
par: Hou, Xianbao, et autres
Publié: (2026)
par: Hou, Xianbao, et autres
Publié: (2026)
Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification
par: Xu, Wenjia, et autres
Publié: (2024)
par: Xu, Wenjia, et autres
Publié: (2024)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
par: Deng, Andong, et autres
Publié: (2024)
par: Deng, Andong, et autres
Publié: (2024)
Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models
par: Zhang, Jielu, et autres
Publié: (2023)
par: Zhang, Jielu, et autres
Publié: (2023)
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
par: Zhou, Qiji, et autres
Publié: (2024)
par: Zhou, Qiji, et autres
Publié: (2024)
Referring Remote Sensing Image Segmentation with Cross-view Semantics Interaction Network
par: Yang, Jiaxing, et autres
Publié: (2025)
par: Yang, Jiaxing, et autres
Publié: (2025)
Learnable Prompt for Few-Shot Semantic Segmentation in Remote Sensing Domain
par: Immanuel, Steve Andreas, et autres
Publié: (2024)
par: Immanuel, Steve Andreas, et autres
Publié: (2024)
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
par: Wang, Fengxiang, et autres
Publié: (2026)
par: Wang, Fengxiang, et autres
Publié: (2026)
Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine
par: Huang, Xiaoshuang, et autres
Publié: (2024)
par: Huang, Xiaoshuang, et autres
Publié: (2024)
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
par: Guo, Grace, et autres
Publié: (2024)
par: Guo, Grace, et autres
Publié: (2024)
Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing
par: Zhang, Weiyu, et autres
Publié: (2026)
par: Zhang, Weiyu, et autres
Publié: (2026)
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
par: Li, Yueying, et autres
Publié: (2026)
par: Li, Yueying, et autres
Publié: (2026)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
par: Özdemir, Övgü, et autres
Publié: (2024)
par: Özdemir, Övgü, et autres
Publié: (2024)
FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing
par: Dang, Yunkai, et autres
Publié: (2025)
par: Dang, Yunkai, et autres
Publié: (2025)
Vision-Language Models in Remote Sensing: Current Progress and Future Trends
par: Li, Xiang, et autres
Publié: (2023)
par: Li, Xiang, et autres
Publié: (2023)
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
par: Xu, Linrui, et autres
Publié: (2024)
par: Xu, Linrui, et autres
Publié: (2024)
Efficient Adaptation For Remote Sensing Visual Grounding
par: Moughnieh, Hasan, et autres
Publié: (2025)
par: Moughnieh, Hasan, et autres
Publié: (2025)
Open-Vocabulary Remote Sensing Image Semantic Segmentation
par: Cao, Qinglong, et autres
Publié: (2024)
par: Cao, Qinglong, et autres
Publié: (2024)
Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation
par: Dong, Zhe, et autres
Publié: (2024)
par: Dong, Zhe, et autres
Publié: (2024)
Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning
par: Xu, Yu, et autres
Publié: (2026)
par: Xu, Yu, et autres
Publié: (2026)
PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
par: Blume, Ansel, et autres
Publié: (2025)
par: Blume, Ansel, et autres
Publié: (2025)
ChangeQuery: Advancing Remote Sensing Change Analysis for Natural and Human-Induced Disasters from Visual Detection to Semantic Understanding
par: Sun, Dongwei, et autres
Publié: (2026)
par: Sun, Dongwei, et autres
Publié: (2026)
Towards Understanding Visual Grounding in Visual Language Models
par: Pantazopoulos, Georgios, et autres
Publié: (2025)
par: Pantazopoulos, Georgios, et autres
Publié: (2025)
HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing
par: Zhang, Xinyu, et autres
Publié: (2026)
par: Zhang, Xinyu, et autres
Publié: (2026)
Documents similaires
-
Aquila: A Hierarchically Aligned Visual-Language Model for Enhanced Remote Sensing Image Comprehension
par: Lu, Kaixuan, et autres
Publié: (2024) -
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
par: Shao, Run, et autres
Publié: (2024) -
SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models
par: Zhong, Chen, et autres
Publié: (2026) -
Beyond Pixels: A Training-Free, Text-to-Text Framework for Remote Sensing Image Retrieval
par: Xiao, J., et autres
Publié: (2025) -
PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval
par: Pan, Jiancheng, et autres
Publié: (2024)