From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Lincan, Kang, Jingxuan, Li, Shuang, Ma, Wenxuan, Xie, Binhui, Qin, Zhida, Liang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation
by: Cai, Lincan, et al.
Published: (2024)
by: Cai, Lincan, et al.
Published: (2024)
Learning Modality Knowledge Alignment for Cross-Modality Transfer
by: Ma, Wenxuan, et al.
Published: (2024)
by: Ma, Wenxuan, et al.
Published: (2024)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
by: Jiang, Yubo, et al.
Published: (2026)
by: Jiang, Yubo, et al.
Published: (2026)
Attention to Detail: Global-Local Attention for High-Resolution AI-Generated Image Detection
by: Han, Lawrence
Published: (2026)
by: Han, Lawrence
Published: (2026)
GLoD: Composing Global Contexts and Local Details in Image Generation
by: Yamada, Moyuru
Published: (2024)
by: Yamada, Moyuru
Published: (2024)
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
by: de Margerie, Anatole Jacquin, et al.
Published: (2025)
by: de Margerie, Anatole Jacquin, et al.
Published: (2025)
Image Super-Resolution Reconstruction Network based on Enhanced Swin Transformer via Alternating Aggregation of Local-Global Features
by: Huang, Yuming, et al.
Published: (2023)
by: Huang, Yuming, et al.
Published: (2023)
Seeing Like Radiologists: Context- and Gaze-Guided Vision-Language Pretraining for Chest X-rays
by: Liu, Kang, et al.
Published: (2026)
by: Liu, Kang, et al.
Published: (2026)
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
by: Xu, Yichen, et al.
Published: (2025)
by: Xu, Yichen, et al.
Published: (2025)
Local-Global Context Aware Transformer for Language-Guided Video Segmentation
by: Liang, Chen, et al.
Published: (2022)
by: Liang, Chen, et al.
Published: (2022)
Local Attention Transformers for High-Detail Optical Flow Upsampling
by: Gielisse, Alexander, et al.
Published: (2024)
by: Gielisse, Alexander, et al.
Published: (2024)
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024)
by: Gou, Chenhui, et al.
Published: (2024)
Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
by: Colagrande, Alex, et al.
Published: (2025)
by: Colagrande, Alex, et al.
Published: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation
by: Li, Boyi, et al.
Published: (2025)
by: Li, Boyi, et al.
Published: (2025)
Leveraging Spatial Attention and Edge Context for Optimized Feature Selection in Visual Localization
by: Istighfarin, Nanda Febri, et al.
Published: (2024)
by: Istighfarin, Nanda Febri, et al.
Published: (2024)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Ramen: Robust Test-Time Adaptation of Vision-Language Models with Active Sample Selection
by: Bao, Wenxuan, et al.
Published: (2026)
by: Bao, Wenxuan, et al.
Published: (2026)
ToSA: Token Selective Attention for Efficient Vision Transformers
by: Singh, Manish Kumar, et al.
Published: (2024)
by: Singh, Manish Kumar, et al.
Published: (2024)
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
by: Tang, Xiaoya, et al.
Published: (2025)
by: Tang, Xiaoya, et al.
Published: (2025)
Generative Dataset Distillation: Balancing Global Structure and Local Details
by: Li, Longzhen, et al.
Published: (2024)
by: Li, Longzhen, et al.
Published: (2024)
Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
by: Liang, Xinyue, et al.
Published: (2025)
by: Liang, Xinyue, et al.
Published: (2025)
Adaptive Global-Local Representation Learning and Selection for Cross-Domain Facial Expression Recognition
by: Gao, Yuefang, et al.
Published: (2024)
by: Gao, Yuefang, et al.
Published: (2024)
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
by: Yuan, Liping, et al.
Published: (2025)
by: Yuan, Liping, et al.
Published: (2025)
Online In-Context Distillation for Low-Resource Vision Language Models
by: Kang, Zhiqi, et al.
Published: (2025)
by: Kang, Zhiqi, et al.
Published: (2025)
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
by: Gan, Chaofan, et al.
Published: (2025)
by: Gan, Chaofan, et al.
Published: (2025)
One Last Attention for Your Vision-Language Model
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
Gaze into the Details: Locality-Sensitive Enhancement for OCTA Retinal Vessel Segmentation
by: Huang, Tuopusen, et al.
Published: (2026)
by: Huang, Tuopusen, et al.
Published: (2026)
Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks
by: Ding, Yuhe, et al.
Published: (2024)
by: Ding, Yuhe, et al.
Published: (2024)
ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
by: Zhang, Yudong, et al.
Published: (2024)
by: Zhang, Yudong, et al.
Published: (2024)
Not All Agents Matter: From Global Attention Dilution to Risk-Prioritized Game Planning
by: Ding, Kang, et al.
Published: (2026)
by: Ding, Kang, et al.
Published: (2026)
Unified Local and Global Attention Interaction Modeling for Vision Transformers
by: Nguyen, Tan, et al.
Published: (2024)
by: Nguyen, Tan, et al.
Published: (2024)
Multi-scale Information Sharing and Selection Network with Boundary Attention for Polyp Segmentation
by: Kang, Xiaolu, et al.
Published: (2024)
by: Kang, Xiaolu, et al.
Published: (2024)
Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations
by: Wang, Yuheng, et al.
Published: (2026)
by: Wang, Yuheng, et al.
Published: (2026)
Global-Local Detail Guided Transformer for Sea Ice Recognition in Optical Remote Sensing Images
by: Huang, Zhanchao, et al.
Published: (2024)
by: Huang, Zhanchao, et al.
Published: (2024)
FlexAttention for Efficient High-Resolution Vision-Language Models
by: Li, Junyan, et al.
Published: (2024)
by: Li, Junyan, et al.
Published: (2024)
Rethinking Causal Mask Attention for Vision-Language Inference
by: Pei, Xiaohuan, et al.
Published: (2025)
by: Pei, Xiaohuan, et al.
Published: (2025)
GalLoP: Learning Global and Local Prompts for Vision-Language Models
by: Lafon, Marc, et al.
Published: (2024)
by: Lafon, Marc, et al.
Published: (2024)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
Similar Items
-
Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation
by: Cai, Lincan, et al.
Published: (2024) -
Learning Modality Knowledge Alignment for Cross-Modality Transfer
by: Ma, Wenxuan, et al.
Published: (2024) -
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
by: Jiang, Yubo, et al.
Published: (2026) -
Attention to Detail: Global-Local Attention for High-Resolution AI-Generated Image Detection
by: Han, Lawrence
Published: (2026) -
GLoD: Composing Global Contexts and Local Details in Image Generation
by: Yamada, Moyuru
Published: (2024)