From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Lincan, Kang, Jingxuan, Li, Shuang, Ma, Wenxuan, Xie, Binhui, Qin, Zhida, Liang, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation
von: Cai, Lincan, et al.
Veröffentlicht: (2024)
von: Cai, Lincan, et al.
Veröffentlicht: (2024)
Learning Modality Knowledge Alignment for Cross-Modality Transfer
von: Ma, Wenxuan, et al.
Veröffentlicht: (2024)
von: Ma, Wenxuan, et al.
Veröffentlicht: (2024)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
von: Jiang, Yubo, et al.
Veröffentlicht: (2026)
von: Jiang, Yubo, et al.
Veröffentlicht: (2026)
Attention to Detail: Global-Local Attention for High-Resolution AI-Generated Image Detection
von: Han, Lawrence
Veröffentlicht: (2026)
von: Han, Lawrence
Veröffentlicht: (2026)
GLoD: Composing Global Contexts and Local Details in Image Generation
von: Yamada, Moyuru
Veröffentlicht: (2024)
von: Yamada, Moyuru
Veröffentlicht: (2024)
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
von: de Margerie, Anatole Jacquin, et al.
Veröffentlicht: (2025)
von: de Margerie, Anatole Jacquin, et al.
Veröffentlicht: (2025)
Image Super-Resolution Reconstruction Network based on Enhanced Swin Transformer via Alternating Aggregation of Local-Global Features
von: Huang, Yuming, et al.
Veröffentlicht: (2023)
von: Huang, Yuming, et al.
Veröffentlicht: (2023)
Seeing Like Radiologists: Context- and Gaze-Guided Vision-Language Pretraining for Chest X-rays
von: Liu, Kang, et al.
Veröffentlicht: (2026)
von: Liu, Kang, et al.
Veröffentlicht: (2026)
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
Local-Global Context Aware Transformer for Language-Guided Video Segmentation
von: Liang, Chen, et al.
Veröffentlicht: (2022)
von: Liang, Chen, et al.
Veröffentlicht: (2022)
Local Attention Transformers for High-Detail Optical Flow Upsampling
von: Gielisse, Alexander, et al.
Veröffentlicht: (2024)
von: Gielisse, Alexander, et al.
Veröffentlicht: (2024)
How Well Can Vision Language Models See Image Details?
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)
Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
von: Colagrande, Alex, et al.
Veröffentlicht: (2025)
von: Colagrande, Alex, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation
von: Li, Boyi, et al.
Veröffentlicht: (2025)
von: Li, Boyi, et al.
Veröffentlicht: (2025)
Leveraging Spatial Attention and Edge Context for Optimized Feature Selection in Visual Localization
von: Istighfarin, Nanda Febri, et al.
Veröffentlicht: (2024)
von: Istighfarin, Nanda Febri, et al.
Veröffentlicht: (2024)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
Ramen: Robust Test-Time Adaptation of Vision-Language Models with Active Sample Selection
von: Bao, Wenxuan, et al.
Veröffentlicht: (2026)
von: Bao, Wenxuan, et al.
Veröffentlicht: (2026)
ToSA: Token Selective Attention for Efficient Vision Transformers
von: Singh, Manish Kumar, et al.
Veröffentlicht: (2024)
von: Singh, Manish Kumar, et al.
Veröffentlicht: (2024)
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
von: Tang, Xiaoya, et al.
Veröffentlicht: (2025)
von: Tang, Xiaoya, et al.
Veröffentlicht: (2025)
Generative Dataset Distillation: Balancing Global Structure and Local Details
von: Li, Longzhen, et al.
Veröffentlicht: (2024)
von: Li, Longzhen, et al.
Veröffentlicht: (2024)
Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
von: Liang, Xinyue, et al.
Veröffentlicht: (2025)
von: Liang, Xinyue, et al.
Veröffentlicht: (2025)
Adaptive Global-Local Representation Learning and Selection for Cross-Domain Facial Expression Recognition
von: Gao, Yuefang, et al.
Veröffentlicht: (2024)
von: Gao, Yuefang, et al.
Veröffentlicht: (2024)
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
von: Yuan, Liping, et al.
Veröffentlicht: (2025)
von: Yuan, Liping, et al.
Veröffentlicht: (2025)
Online In-Context Distillation for Low-Resource Vision Language Models
von: Kang, Zhiqi, et al.
Veröffentlicht: (2025)
von: Kang, Zhiqi, et al.
Veröffentlicht: (2025)
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)
One Last Attention for Your Vision-Language Model
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Gaze into the Details: Locality-Sensitive Enhancement for OCTA Retinal Vessel Segmentation
von: Huang, Tuopusen, et al.
Veröffentlicht: (2026)
von: Huang, Tuopusen, et al.
Veröffentlicht: (2026)
Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks
von: Ding, Yuhe, et al.
Veröffentlicht: (2024)
von: Ding, Yuhe, et al.
Veröffentlicht: (2024)
ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
Not All Agents Matter: From Global Attention Dilution to Risk-Prioritized Game Planning
von: Ding, Kang, et al.
Veröffentlicht: (2026)
von: Ding, Kang, et al.
Veröffentlicht: (2026)
Unified Local and Global Attention Interaction Modeling for Vision Transformers
von: Nguyen, Tan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan, et al.
Veröffentlicht: (2024)
Multi-scale Information Sharing and Selection Network with Boundary Attention for Polyp Segmentation
von: Kang, Xiaolu, et al.
Veröffentlicht: (2024)
von: Kang, Xiaolu, et al.
Veröffentlicht: (2024)
Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations
von: Wang, Yuheng, et al.
Veröffentlicht: (2026)
von: Wang, Yuheng, et al.
Veröffentlicht: (2026)
Global-Local Detail Guided Transformer for Sea Ice Recognition in Optical Remote Sensing Images
von: Huang, Zhanchao, et al.
Veröffentlicht: (2024)
von: Huang, Zhanchao, et al.
Veröffentlicht: (2024)
FlexAttention for Efficient High-Resolution Vision-Language Models
von: Li, Junyan, et al.
Veröffentlicht: (2024)
von: Li, Junyan, et al.
Veröffentlicht: (2024)
Rethinking Causal Mask Attention for Vision-Language Inference
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2025)
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2025)
GalLoP: Learning Global and Local Prompts for Vision-Language Models
von: Lafon, Marc, et al.
Veröffentlicht: (2024)
von: Lafon, Marc, et al.
Veröffentlicht: (2024)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
von: Huang, Jen-tse, et al.
Veröffentlicht: (2025)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation
von: Cai, Lincan, et al.
Veröffentlicht: (2024) -
Learning Modality Knowledge Alignment for Cross-Modality Transfer
von: Ma, Wenxuan, et al.
Veröffentlicht: (2024) -
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
von: Jiang, Yubo, et al.
Veröffentlicht: (2026) -
Attention to Detail: Global-Local Attention for High-Resolution AI-Generated Image Detection
von: Han, Lawrence
Veröffentlicht: (2026) -
GLoD: Composing Global Contexts and Local Details in Image Generation
von: Yamada, Moyuru
Veröffentlicht: (2024)