Beyond Dominant Patches: Spatial Credit Redistribution For Grounded Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Samin, Niamul Hassan, Rahman, Md Arifur, Arean, Abdullah Ibne Hanif, Noshin, Juena Ahmed, Rahman, Md Ashikur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
Automatic Vehicle Detection using DETR: A Transformer-Based Approach for Navigating Treacherous Roads
von: Fahad, Istiaq Ahmed, et al.
Veröffentlicht: (2025)
von: Fahad, Istiaq Ahmed, et al.
Veröffentlicht: (2025)
BanglaMM-Disaster: A Multimodal Transformer-Based Deep Learning Framework for Multiclass Disaster Classification in Bangla
von: Islam, Ariful, et al.
Veröffentlicht: (2025)
von: Islam, Ariful, et al.
Veröffentlicht: (2025)
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
von: Chowdhury, Md Abtahi Majeed, et al.
Veröffentlicht: (2025)
von: Chowdhury, Md Abtahi Majeed, et al.
Veröffentlicht: (2025)
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models
von: Monon, Mashrafi, et al.
Veröffentlicht: (2026)
von: Monon, Mashrafi, et al.
Veröffentlicht: (2026)
Two Decades of Bengali Handwritten Digit Recognition: A Survey
von: Rahman, A. B. M. Ashikur, et al.
Veröffentlicht: (2022)
von: Rahman, A. B. M. Ashikur, et al.
Veröffentlicht: (2022)
ColorFoil: Investigating Color Blindness in Large Vision and Language Models
von: Samin, Ahnaf Mozib, et al.
Veröffentlicht: (2024)
von: Samin, Ahnaf Mozib, et al.
Veröffentlicht: (2024)
BeHGAN: Bengali Handwritten Word Generation from Plain Text Using Generative Adversarial Networks
von: Islam, Md. Rakibul, et al.
Veröffentlicht: (2025)
von: Islam, Md. Rakibul, et al.
Veröffentlicht: (2025)
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
MosquitoFusion: A Multiclass Dataset for Real-Time Detection of Mosquitoes, Swarms, and Breeding Sites Using Deep Learning
von: Sayeedi, Md. Faiyaz Abdullah, et al.
Veröffentlicht: (2024)
von: Sayeedi, Md. Faiyaz Abdullah, et al.
Veröffentlicht: (2024)
CAST: Channel-Aware Spatial Transfer Learning with Pseudo-Image Radar for Sign Language Recognition
von: Shujon, Md. Shakhoyat Rahman, et al.
Veröffentlicht: (2026)
von: Shujon, Md. Shakhoyat Rahman, et al.
Veröffentlicht: (2026)
In Pursuit of Many: A Review of Modern Multiple Object Tracking Systems
von: Bashar, Mk, et al.
Veröffentlicht: (2022)
von: Bashar, Mk, et al.
Veröffentlicht: (2022)
OncoVision: Integrating Mammography and Clinical Data through Attention-Driven Multimodal AI for Enhanced Breast Cancer Diagnosis
von: Ahmed, Istiak, et al.
Veröffentlicht: (2025)
von: Ahmed, Istiak, et al.
Veröffentlicht: (2025)
Pattern Recognition Tasks with Personalized Federated Learning
von: Rahman, Md. Arifur, et al.
Veröffentlicht: (2026)
von: Rahman, Md. Arifur, et al.
Veröffentlicht: (2026)
BdSL-SPOTER: A Transformer-Based Framework for Bengali Sign Language Recognition with Cultural Adaptation
von: Azad, Sayad Ibna, et al.
Veröffentlicht: (2025)
von: Azad, Sayad Ibna, et al.
Veröffentlicht: (2025)
Comparative Performance Analysis of Transformer-Based Pre-Trained Models for Detecting Keratoconus Disease
von: Ahmed, Nayeem, et al.
Veröffentlicht: (2024)
von: Ahmed, Nayeem, et al.
Veröffentlicht: (2024)
EVCC: Enhanced Vision Transformer-ConvNeXt-CoAtNet Fusion for Classification
von: Hasan, Kazi Reyazul, et al.
Veröffentlicht: (2025)
von: Hasan, Kazi Reyazul, et al.
Veröffentlicht: (2025)
MK-UNet: Multi-kernel Lightweight CNN for Medical Image Segmentation
von: Rahman, Md Mostafijur, et al.
Veröffentlicht: (2025)
von: Rahman, Md Mostafijur, et al.
Veröffentlicht: (2025)
LoMix: Learnable Weighted Multi-Scale Logits Mixing for Medical Image Segmentation
von: Rahman, Md Mostafijur, et al.
Veröffentlicht: (2025)
von: Rahman, Md Mostafijur, et al.
Veröffentlicht: (2025)
A Heterogeneous Two-Stream Framework for Video Action Recognition with Comparative Fusion Analysis
von: Rahaman, Md. Afzalur, et al.
Veröffentlicht: (2026)
von: Rahaman, Md. Afzalur, et al.
Veröffentlicht: (2026)
IMVB7t: A Multi-Modal Model for Food Preferences based on Artificially Produced Traits
von: Abir, Mushfiqur Rahman, et al.
Veröffentlicht: (2024)
von: Abir, Mushfiqur Rahman, et al.
Veröffentlicht: (2024)
To Agree or To Be Right? The Grounding-Sycophancy Tradeoff in Medical Vision-Language Models
von: Aranya, OFM Riaz Rahman, et al.
Veröffentlicht: (2026)
von: Aranya, OFM Riaz Rahman, et al.
Veröffentlicht: (2026)
Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models
von: Alam, Md Tanvirul
Veröffentlicht: (2026)
von: Alam, Md Tanvirul
Veröffentlicht: (2026)
AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
von: Munir, Mustafa, et al.
Veröffentlicht: (2025)
von: Munir, Mustafa, et al.
Veröffentlicht: (2025)
Multi-hop Relational Contrastive Learning: Extending Spatial Contrastive Pre-training Beyond Pairwise Relations
von: Ahmed, Sheikh Tanvir, et al.
Veröffentlicht: (2026)
von: Ahmed, Sheikh Tanvir, et al.
Veröffentlicht: (2026)
Adaptive Enhancement and Dual-Pooling Sequential Attention for Lightweight Underwater Object Detection with YOLOv10
von: Rahman, Md. Mushibur, et al.
Veröffentlicht: (2026)
von: Rahman, Md. Mushibur, et al.
Veröffentlicht: (2026)
VisionTrap: Unanswerable Questions On Visual Data
von: Saadat, Asir, et al.
Veröffentlicht: (2025)
von: Saadat, Asir, et al.
Veröffentlicht: (2025)
BrainRotViT: Transformer-ResNet Hybrid for Explainable Modeling of Brain Aging from 3D sMRI
von: Jalal, Wasif, et al.
Veröffentlicht: (2025)
von: Jalal, Wasif, et al.
Veröffentlicht: (2025)
BanglaNet: Bangla Handwritten Character Recognition using Ensembling of Convolutional Neural Network
von: Saha, Chandrika, et al.
Veröffentlicht: (2024)
von: Saha, Chandrika, et al.
Veröffentlicht: (2024)
IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2025)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2025)
Spatial Transcriptomics Analysis of Zero-shot Gene Expression Prediction
von: Yang, Yan, et al.
Veröffentlicht: (2024)
von: Yang, Yan, et al.
Veröffentlicht: (2024)
VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites
von: Islam, Md. Adnanul, et al.
Veröffentlicht: (2025)
von: Islam, Md. Adnanul, et al.
Veröffentlicht: (2025)
Probabilistic Feature Imputation and Uncertainty-Aware Multimodal Federated Aggregation
von: Shahid, Nafis Fuad, et al.
Veröffentlicht: (2026)
von: Shahid, Nafis Fuad, et al.
Veröffentlicht: (2026)
Aerial Flood Scene Classification Using Fine-Tuned Attention-based Architecture for Flood-Prone Countries in South Asia
von: Hassan, Ibne, et al.
Veröffentlicht: (2024)
von: Hassan, Ibne, et al.
Veröffentlicht: (2024)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection
von: Abdullah, Ahmed, et al.
Veröffentlicht: (2026)
von: Abdullah, Ahmed, et al.
Veröffentlicht: (2026)
ResNet-34 with Lightweight Decoder for Accurate and Efficient Segmentation of Fetal Brain MRI
von: Rahman, Ashiqur, et al.
Veröffentlicht: (2026)
von: Rahman, Ashiqur, et al.
Veröffentlicht: (2026)
A Two-Stage Deep Learning Framework for Segmentation of Ten Gastrointestinal Organs from Coronal MR Enterography
von: Rahman, Ashiqur, et al.
Veröffentlicht: (2026)
von: Rahman, Ashiqur, et al.
Veröffentlicht: (2026)
An empirical study for the early detection of Mpox from skin lesion images using pretrained CNN models leveraging XAI technique
von: Rahim, Mohammad Asifur, et al.
Veröffentlicht: (2025)
von: Rahim, Mohammad Asifur, et al.
Veröffentlicht: (2025)
Vision Transformers for End-to-End Quark-Gluon Jet Classification from Calorimeter Images
von: Jahin, Md Abrar, et al.
Veröffentlicht: (2025)
von: Jahin, Md Abrar, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026) -
Automatic Vehicle Detection using DETR: A Transformer-Based Approach for Navigating Treacherous Roads
von: Fahad, Istiaq Ahmed, et al.
Veröffentlicht: (2025) -
BanglaMM-Disaster: A Multimodal Transformer-Based Deep Learning Framework for Multiclass Disaster Classification in Bangla
von: Islam, Ariful, et al.
Veröffentlicht: (2025) -
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
von: Chowdhury, Md Abtahi Majeed, et al.
Veröffentlicht: (2025) -
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models
von: Monon, Mashrafi, et al.
Veröffentlicht: (2026)