Head-Aware Visual Cropping: Enhancing Fine-Grained VQA with Attention-Guided Subimage
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Junfei, Pan, Peng, Zhang, Xulong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
by: Ma, Yiming, et al.
Published: (2026)
by: Ma, Yiming, et al.
Published: (2026)
Multi-Perspective Subimage CLIP with Keyword Guidance for Remote Sensing Image-Text Retrieval
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023)
by: Zhang, Xiaoman, et al.
Published: (2023)
Multispectral Fine-Grained Classification of Blackgrass in Wheat and Barley Crops
by: Darbyshire, Madeleine, et al.
Published: (2024)
by: Darbyshire, Madeleine, et al.
Published: (2024)
Visually-Guided Controllable Medical Image Generation via Fine-Grained Semantic Disentanglement
by: Huang, Xin, et al.
Published: (2026)
by: Huang, Xin, et al.
Published: (2026)
Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation
by: Lian, Sheng, et al.
Published: (2025)
by: Lian, Sheng, et al.
Published: (2025)
TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
by: Zhang, Zijian, et al.
Published: (2025)
by: Zhang, Zijian, et al.
Published: (2025)
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
by: Diao, Xingjian, et al.
Published: (2025)
by: Diao, Xingjian, et al.
Published: (2025)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
Context-Semantic Quality Awareness Network for Fine-Grained Visual Categorization
by: Xu, Qin, et al.
Published: (2024)
by: Xu, Qin, et al.
Published: (2024)
DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
by: Li, Geng, et al.
Published: (2025)
by: Li, Geng, et al.
Published: (2025)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
On the Role of Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2024)
by: Reich, Daniel, et al.
Published: (2024)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
Toward Fine-Grained Facial Control in 3D Talking Head Generation
by: Xie, Shaoyang, et al.
Published: (2026)
by: Xie, Shaoyang, et al.
Published: (2026)
Uncertainty Guided Refinement for Fine-Grained Salient Object Detection
by: Yuan, Yao, et al.
Published: (2025)
by: Yuan, Yao, et al.
Published: (2025)
HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models
by: Sayem, MD Khalequzzaman Chowdhury, et al.
Published: (2026)
by: Sayem, MD Khalequzzaman Chowdhury, et al.
Published: (2026)
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
by: Tateno, Masatoshi, et al.
Published: (2025)
by: Tateno, Masatoshi, et al.
Published: (2025)
From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception
by: Zhu, Jilong, et al.
Published: (2026)
by: Zhu, Jilong, et al.
Published: (2026)
HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models
by: Zhu, Qihui, et al.
Published: (2026)
by: Zhu, Qihui, et al.
Published: (2026)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection
by: Wang, Hui, et al.
Published: (2026)
by: Wang, Hui, et al.
Published: (2026)
RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding
by: Sun, Muyi, et al.
Published: (2026)
by: Sun, Muyi, et al.
Published: (2026)
Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language Models
by: Hai, Jia-Wei, et al.
Published: (2026)
by: Hai, Jia-Wei, et al.
Published: (2026)
SGIA: Enhancing Fine-Grained Visual Classification with Sequence Generative Image Augmentation
by: Liao, Qiyu, et al.
Published: (2024)
by: Liao, Qiyu, et al.
Published: (2024)
Synchronized and Fine-Grained Head for Skeleton-Based Ambiguous Action Recognition
by: Huang, Hao, et al.
Published: (2024)
by: Huang, Hao, et al.
Published: (2024)
Measuring Faithful and Plausible Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2023)
by: Reich, Daniel, et al.
Published: (2023)
HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps
by: Zhong, Xuchang, et al.
Published: (2026)
by: Zhong, Xuchang, et al.
Published: (2026)
Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach
by: Bi, Jing, et al.
Published: (2024)
by: Bi, Jing, et al.
Published: (2024)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
by: Sood, Ekta, et al.
Published: (2021)
by: Sood, Ekta, et al.
Published: (2021)
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
by: Shi, Jiaqi, et al.
Published: (2026)
by: Shi, Jiaqi, et al.
Published: (2026)
AesCrop: Aesthetic-driven Cropping Guided by Composition
by: Wong, Yen-Hong, et al.
Published: (2025)
by: Wong, Yen-Hong, et al.
Published: (2025)
QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems
by: He, Zhixian, et al.
Published: (2024)
by: He, Zhixian, et al.
Published: (2024)
FG-CLIP: Fine-Grained Visual and Textual Alignment
by: Xie, Chunyu, et al.
Published: (2025)
by: Xie, Chunyu, et al.
Published: (2025)
Integrating Query-aware Segmentation and Cross-Attention for Robust VQA
by: Choi, Wonjun, et al.
Published: (2024)
by: Choi, Wonjun, et al.
Published: (2024)
MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species
by: Lee, Donghwan, et al.
Published: (2026)
by: Lee, Donghwan, et al.
Published: (2026)
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
by: Li, Jiaye, et al.
Published: (2025)
by: Li, Jiaye, et al.
Published: (2025)
FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention
by: Zhang, Peng, et al.
Published: (2025)
by: Zhang, Peng, et al.
Published: (2025)
Enhancing Open-Vocabulary Object Detection through Multi-Level Fine-Grained Visual-Language Alignment
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
Similar Items
-
DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
by: Ma, Yiming, et al.
Published: (2026) -
Multi-Perspective Subimage CLIP with Keyword Guidance for Remote Sensing Image-Text Retrieval
by: Li, Yifan, et al.
Published: (2026) -
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023) -
Multispectral Fine-Grained Classification of Blackgrass in Wheat and Barley Crops
by: Darbyshire, Madeleine, et al.
Published: (2024) -
Visually-Guided Controllable Medical Image Generation via Fine-Grained Semantic Disentanglement
by: Huang, Xin, et al.
Published: (2026)