Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jung, Mingi, Lee, Saehyung, Kim, Eunji, Yoon, Sungroh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
von: Lew, Jaihyun, et al.
Veröffentlicht: (2024)
von: Lew, Jaihyun, et al.
Veröffentlicht: (2024)
On mitigating stability-plasticity dilemma in CLIP-guided image morphing via geodesic distillation loss
von: Oh, Yeongtak, et al.
Veröffentlicht: (2024)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2024)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
ReflectCAP: Detailed Image Captioning with Reflective Memory
von: Min, Kyungmin, et al.
Veröffentlicht: (2026)
von: Min, Kyungmin, et al.
Veröffentlicht: (2026)
Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors
von: Lee, Jonghyun, et al.
Veröffentlicht: (2024)
von: Lee, Jonghyun, et al.
Veröffentlicht: (2024)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
von: Park, Sangha, et al.
Veröffentlicht: (2025)
von: Park, Sangha, et al.
Veröffentlicht: (2025)
Generating Accurate and Detailed Captions for High-Resolution Images
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
Normality Addition via Normality Detection in Industrial Image Anomaly Detection Models
von: Yi, Jihun, et al.
Veröffentlicht: (2024)
von: Yi, Jihun, et al.
Veröffentlicht: (2024)
DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
von: Shin, Chaehun, et al.
Veröffentlicht: (2024)
von: Shin, Chaehun, et al.
Veröffentlicht: (2024)
STAG: Structural Test-time Alignment of Gradients for Online Adaptation
von: Shin, Juhyeon, et al.
Veröffentlicht: (2024)
von: Shin, Juhyeon, et al.
Veröffentlicht: (2024)
Bi-directional Contextual Attention for 3D Dense Captioning
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
Crafting Query-Aware Selective Attention for Single Image Super-Resolution
von: Kim, Junyoung, et al.
Veröffentlicht: (2025)
von: Kim, Junyoung, et al.
Veröffentlicht: (2025)
Attention Via Convolutional Nearest Neighbors
von: Kang, Mingi, et al.
Veröffentlicht: (2025)
von: Kang, Mingi, et al.
Veröffentlicht: (2025)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
von: Jung, Woojun, et al.
Veröffentlicht: (2025)
von: Jung, Woojun, et al.
Veröffentlicht: (2025)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
von: Teja, L. D. M. S. Sai, et al.
Veröffentlicht: (2025)
von: Teja, L. D. M. S. Sai, et al.
Veröffentlicht: (2025)
Benchmarking and Improving Detail Image Caption
von: Dong, Hongyuan, et al.
Veröffentlicht: (2024)
von: Dong, Hongyuan, et al.
Veröffentlicht: (2024)
Visual Representation Alignment for Multimodal Large Language Models
von: Yoon, Heeji, et al.
Veröffentlicht: (2025)
von: Yoon, Heeji, et al.
Veröffentlicht: (2025)
Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization
von: Song, Sanghyeob, et al.
Veröffentlicht: (2024)
von: Song, Sanghyeob, et al.
Veröffentlicht: (2024)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
von: Li, Qiming, et al.
Veröffentlicht: (2026)
von: Li, Qiming, et al.
Veröffentlicht: (2026)
Attention to Detail: Global-Local Attention for High-Resolution AI-Generated Image Detection
von: Han, Lawrence
Veröffentlicht: (2026)
von: Han, Lawrence
Veröffentlicht: (2026)
Selective LoRA for Visual Tokens and Attention Heads
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
von: Baek, Kanghyun, et al.
Veröffentlicht: (2026)
von: Baek, Kanghyun, et al.
Veröffentlicht: (2026)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
von: Cai, Lincan, et al.
Veröffentlicht: (2025)
von: Cai, Lincan, et al.
Veröffentlicht: (2025)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
von: Wang, Pengfei, et al.
Veröffentlicht: (2025)
von: Wang, Pengfei, et al.
Veröffentlicht: (2025)
CKNN: Cleansed k-Nearest Neighbor for Unsupervised Video Anomaly Detection
von: Yi, Jihun, et al.
Veröffentlicht: (2024)
von: Yi, Jihun, et al.
Veröffentlicht: (2024)
TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
von: Baek, Kanghyun, et al.
Veröffentlicht: (2025)
von: Baek, Kanghyun, et al.
Veröffentlicht: (2025)
An Ensemble Model with Attention Based Mechanism for Image Captioning
von: Badarneh, Israa Al, et al.
Veröffentlicht: (2025)
von: Badarneh, Israa Al, et al.
Veröffentlicht: (2025)
Physics-Grounded Adversarial Stain Augmentation with Calibrated Coverage Guarantees
von: Hong, Mingi
Veröffentlicht: (2026)
von: Hong, Mingi
Veröffentlicht: (2026)
Guided Attention for Interpretable Motion Captioning
von: Radouane, Karim, et al.
Veröffentlicht: (2023)
von: Radouane, Karim, et al.
Veröffentlicht: (2023)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
von: Song, Zijie, et al.
Veröffentlicht: (2023)
von: Song, Zijie, et al.
Veröffentlicht: (2023)
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
von: Li, Qiming, et al.
Veröffentlicht: (2025)
von: Li, Qiming, et al.
Veröffentlicht: (2025)
Attention Calibration for Disentangled Text-to-Image Personalization
von: Zhang, Yanbing, et al.
Veröffentlicht: (2024)
von: Zhang, Yanbing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
von: Lee, Saehyung, et al.
Veröffentlicht: (2024) -
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
von: Lee, Saehyung, et al.
Veröffentlicht: (2024) -
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
von: Lee, Saehyung, et al.
Veröffentlicht: (2024) -
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
von: Lew, Jaihyun, et al.
Veröffentlicht: (2024) -
On mitigating stability-plasticity dilemma in CLIP-guided image morphing via geodesic distillation loss
von: Oh, Yeongtak, et al.
Veröffentlicht: (2024)