Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Sujung, Yoon, Chanyong, Hwang, Seong Jae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DragText: Rethinking Text Embedding in Point-based Image Editing
von: Choi, Gayoon, et al.
Veröffentlicht: (2024)
von: Choi, Gayoon, et al.
Veröffentlicht: (2024)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
von: Song, Chull Hwan, et al.
Veröffentlicht: (2024)
von: Song, Chull Hwan, et al.
Veröffentlicht: (2024)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
Large Language Models Facilitate Vision Reflection in Image Classification
von: An, Guoyuan, et al.
Veröffentlicht: (2025)
von: An, Guoyuan, et al.
Veröffentlicht: (2025)
FALCON: Frequency Adjoint Link with CONtinuous Density Mask for Fast Single Image Dehazing
von: Kim, Donghyun, et al.
Veröffentlicht: (2024)
von: Kim, Donghyun, et al.
Veröffentlicht: (2024)
Mono-Modalizing Extremely Heterogeneous Multi-Modal Medical Image Registration
von: Choo, Kyobin, et al.
Veröffentlicht: (2025)
von: Choo, Kyobin, et al.
Veröffentlicht: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
IKOD: Mitigating Visual Attention Degradation in Large Vision-Language Models
von: Yang, Jiabing, et al.
Veröffentlicht: (2025)
von: Yang, Jiabing, et al.
Veröffentlicht: (2025)
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
von: Zhou, Guanyu, et al.
Veröffentlicht: (2024)
von: Zhou, Guanyu, et al.
Veröffentlicht: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
Learning Context-Adaptive Motion Priors for Masked Motion Diffusion Models with Efficient Kinematic Attention Aggregation
von: Jiang, Junkun, et al.
Veröffentlicht: (2026)
von: Jiang, Junkun, et al.
Veröffentlicht: (2026)
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
von: Yang, Chengxu, et al.
Veröffentlicht: (2026)
von: Yang, Chengxu, et al.
Veröffentlicht: (2026)
Masked Diffusion Vision-Language Models for Temporal Action Localization
von: Wang, Fengshun, et al.
Veröffentlicht: (2026)
von: Wang, Fengshun, et al.
Veröffentlicht: (2026)
Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
von: Li, Qiming, et al.
Veröffentlicht: (2025)
von: Li, Qiming, et al.
Veröffentlicht: (2025)
Self-Correction Inside the Model: Leveraging Layer Attention to Mitigate Hallucinations in Large Vision Language Models
von: Fu, April
Veröffentlicht: (2026)
von: Fu, April
Veröffentlicht: (2026)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025)
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025)
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
von: Tian, Xinyu, et al.
Veröffentlicht: (2025)
von: Tian, Xinyu, et al.
Veröffentlicht: (2025)
Beyond Semantic Priors: Mitigating Optimization Collapse for Generalizable Visual Forensics
von: Liu, Jipeng, et al.
Veröffentlicht: (2026)
von: Liu, Jipeng, et al.
Veröffentlicht: (2026)
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025)
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
von: Han, Woojung, et al.
Veröffentlicht: (2025)
von: Han, Woojung, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
von: Zhu, Younan, et al.
Veröffentlicht: (2025)
von: Zhu, Younan, et al.
Veröffentlicht: (2025)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
von: Song, Jiale, et al.
Veröffentlicht: (2026)
von: Song, Jiale, et al.
Veröffentlicht: (2026)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
von: Li, Qiming, et al.
Veröffentlicht: (2026)
von: Li, Qiming, et al.
Veröffentlicht: (2026)
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026)
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026)
LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models
von: Wu, Zongyu, et al.
Veröffentlicht: (2025)
von: Wu, Zongyu, et al.
Veröffentlicht: (2025)
Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain
von: Lee, Seulbi, et al.
Veröffentlicht: (2026)
von: Lee, Seulbi, et al.
Veröffentlicht: (2026)
MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
von: Li, Chenxi, et al.
Veröffentlicht: (2025)
von: Li, Chenxi, et al.
Veröffentlicht: (2025)
Exploring Causes and Mitigation of Hallucinations in Large Vision Language Models
von: Sun, Yaqi, et al.
Veröffentlicht: (2025)
von: Sun, Yaqi, et al.
Veröffentlicht: (2025)
ViRAC: A Vision-Reasoning Agent Head Movement Control Framework in Arbitrary Virtual Environments
von: Hwang, Juyeong, et al.
Veröffentlicht: (2025)
von: Hwang, Juyeong, et al.
Veröffentlicht: (2025)
Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
von: Jiang, Zhangqi, et al.
Veröffentlicht: (2024)
von: Jiang, Zhangqi, et al.
Veröffentlicht: (2024)
Polyline Path Masked Attention for Vision Transformer
von: Zhao, Zhongchen, et al.
Veröffentlicht: (2025)
von: Zhao, Zhongchen, et al.
Veröffentlicht: (2025)
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
von: Sun, Yiming, et al.
Veröffentlicht: (2025)
von: Sun, Yiming, et al.
Veröffentlicht: (2025)
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
von: Ren, Lingfeng, et al.
Veröffentlicht: (2026)
von: Ren, Lingfeng, et al.
Veröffentlicht: (2026)
Rethinking Causal Mask Attention for Vision-Language Inference
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2025)
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2025)
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
von: Jeong, Taejin, et al.
Veröffentlicht: (2026)
von: Jeong, Taejin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DragText: Rethinking Text Embedding in Point-based Image Editing
von: Choi, Gayoon, et al.
Veröffentlicht: (2024) -
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
von: Kang, Seil, et al.
Veröffentlicht: (2025) -
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
von: Song, Chull Hwan, et al.
Veröffentlicht: (2024) -
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025) -
Large Language Models Facilitate Vision Reflection in Image Classification
von: An, Guoyuan, et al.
Veröffentlicht: (2025)