Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kang, Seil, Kim, Jinyeong, Kim, Junhyeok, Hwang, Seong Jae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
See What You Are Told: Visual Attention Sink in Large Multimodal Models
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
Real-Time Visual Attribution Streaming in Thinking Model
von: Kang, Seil, et al.
Veröffentlicht: (2026)
von: Kang, Seil, et al.
Veröffentlicht: (2026)
Interpreting vision transformers via residual replacement model
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
von: Jeong, Taejin, et al.
Veröffentlicht: (2026)
von: Jeong, Taejin, et al.
Veröffentlicht: (2026)
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026)
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
von: Jun, Youngjun, et al.
Veröffentlicht: (2026)
von: Jun, Youngjun, et al.
Veröffentlicht: (2026)
FALCON: Frequency Adjoint Link with CONtinuous Density Mask for Fast Single Image Dehazing
von: Kim, Donghyun, et al.
Veröffentlicht: (2024)
von: Kim, Donghyun, et al.
Veröffentlicht: (2024)
WoLF: Wide-scope Large Language Model Framework for CXR Understanding
von: Kang, Seil, et al.
Veröffentlicht: (2024)
von: Kang, Seil, et al.
Veröffentlicht: (2024)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2026)
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2026)
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
von: Lee, Dong-Jae, et al.
Veröffentlicht: (2026)
von: Lee, Dong-Jae, et al.
Veröffentlicht: (2026)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
von: Kim, Keuntae, et al.
Veröffentlicht: (2026)
von: Kim, Keuntae, et al.
Veröffentlicht: (2026)
Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
von: Kim, Eunki, et al.
Veröffentlicht: (2025)
von: Kim, Eunki, et al.
Veröffentlicht: (2025)
Selective LoRA for Visual Tokens and Attention Heads
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
von: Kim, Sohee, et al.
Veröffentlicht: (2025)
von: Kim, Sohee, et al.
Veröffentlicht: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
von: An, Na Min, et al.
Veröffentlicht: (2025)
von: An, Na Min, et al.
Veröffentlicht: (2025)
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
von: Kang, Beomseok, et al.
Veröffentlicht: (2026)
von: Kang, Beomseok, et al.
Veröffentlicht: (2026)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
Re-Scoring Using Image-Language Similarity for Few-Shot Object Detection
von: Jung, Min Jae, et al.
Veröffentlicht: (2023)
von: Jung, Min Jae, et al.
Veröffentlicht: (2023)
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
von: Kim, Jihyo, et al.
Veröffentlicht: (2024)
von: Kim, Jihyo, et al.
Veröffentlicht: (2024)
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models
von: Hong, Sujung, et al.
Veröffentlicht: (2026)
von: Hong, Sujung, et al.
Veröffentlicht: (2026)
QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2025)
von: Zhang, Yudong, et al.
Veröffentlicht: (2025)
VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2025)
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2025)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
von: Yan, Siming, et al.
Veröffentlicht: (2024)
von: Yan, Siming, et al.
Veröffentlicht: (2024)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
von: Kim, Mingyeong, et al.
Veröffentlicht: (2026)
von: Kim, Mingyeong, et al.
Veröffentlicht: (2026)
Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
von: Ernhofer, Benjamin Raphael, et al.
Veröffentlicht: (2025)
von: Ernhofer, Benjamin Raphael, et al.
Veröffentlicht: (2025)
Large Vision-Language Models Get Lost in Attention
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
Attention Prompting on Image for Large Vision-Language Models
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
Resource-Efficient Medical Report Generation using Large Language Models
von: Abdullah, et al.
Veröffentlicht: (2024)
von: Abdullah, et al.
Veröffentlicht: (2024)
CoBra: Complementary Branch Fusing Class and Semantic Knowledge for Robust Weakly Supervised Semantic Segmentation
von: Han, Woojung, et al.
Veröffentlicht: (2024)
von: Han, Woojung, et al.
Veröffentlicht: (2024)
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
von: Kang, Seongjae, et al.
Veröffentlicht: (2025)
von: Kang, Seongjae, et al.
Veröffentlicht: (2025)
Language-Guided Invariance Probing of Vision-Language Models
von: Lee, Jae Joong
Veröffentlicht: (2025)
von: Lee, Jae Joong
Veröffentlicht: (2025)
Do We Really Need a Large Number of Visual Prompts?
von: Kim, Youngeun, et al.
Veröffentlicht: (2023)
von: Kim, Youngeun, et al.
Veröffentlicht: (2023)
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
von: Yu, Lu, et al.
Veröffentlicht: (2024)
von: Yu, Lu, et al.
Veröffentlicht: (2024)
PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models
von: Meng, Yu, et al.
Veröffentlicht: (2025)
von: Meng, Yu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
See What You Are Told: Visual Attention Sink in Large Multimodal Models
von: Kang, Seil, et al.
Veröffentlicht: (2025) -
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025) -
Real-Time Visual Attribution Streaming in Thinking Model
von: Kang, Seil, et al.
Veröffentlicht: (2026) -
Interpreting vision transformers via residual replacement model
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025) -
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
von: Jeong, Taejin, et al.
Veröffentlicht: (2026)