Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Jiale, Luo, Jiaxin, Tang, Xue-song, Hao, Kuangrong, Zhao, Mingbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025)
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025)
Mitigating Image Captioning Hallucinations in Vision-Language Models
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
von: Fan, Linfeng, et al.
Veröffentlicht: (2026)
von: Fan, Linfeng, et al.
Veröffentlicht: (2026)
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
von: Chen, Yangneng, et al.
Veröffentlicht: (2026)
von: Chen, Yangneng, et al.
Veröffentlicht: (2026)
Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models
von: Tang, Hao, et al.
Veröffentlicht: (2026)
von: Tang, Hao, et al.
Veröffentlicht: (2026)
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
Multi-Target Federated Backdoor Attack Based on Feature Aggregation
von: Hao, Lingguag, et al.
Veröffentlicht: (2025)
von: Hao, Lingguag, et al.
Veröffentlicht: (2025)
Failures to Surface Harmful Contents in Video Large Language Models
von: Cao, Yuxin, et al.
Veröffentlicht: (2025)
von: Cao, Yuxin, et al.
Veröffentlicht: (2025)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
von: Deng, Ailin, et al.
Veröffentlicht: (2024)
von: Deng, Ailin, et al.
Veröffentlicht: (2024)
Linguistics-Vision Monotonic Consistent Network for Sign Language Production
von: Wang, Xu, et al.
Veröffentlicht: (2024)
von: Wang, Xu, et al.
Veröffentlicht: (2024)
Multi-Scale and Detail-Enhanced Segment Anything Model for Salient Object Detection
von: Gao, Shixuan, et al.
Veröffentlicht: (2024)
von: Gao, Shixuan, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
VAAS: Vision-Attention Anomaly Scoring for Image Manipulation Detection in Digital Forensics
von: Bamigbade, Opeyemi, et al.
Veröffentlicht: (2025)
von: Bamigbade, Opeyemi, et al.
Veröffentlicht: (2025)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
von: Ding, Peng, et al.
Veröffentlicht: (2024)
von: Ding, Peng, et al.
Veröffentlicht: (2024)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
von: Chen, Rongjun, et al.
Veröffentlicht: (2025)
von: Chen, Rongjun, et al.
Veröffentlicht: (2025)
KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object Detection
von: Li, Xingyuan, et al.
Veröffentlicht: (2025)
von: Li, Xingyuan, et al.
Veröffentlicht: (2025)
On the Robustness of Human-Object Interaction Detection against Distribution Shift
von: Xie, Chi, et al.
Veröffentlicht: (2025)
von: Xie, Chi, et al.
Veröffentlicht: (2025)
MSCT: Differential Cross-Modal Attention for Deepfake Detection
von: Wei, Fangda, et al.
Veröffentlicht: (2026)
von: Wei, Fangda, et al.
Veröffentlicht: (2026)
Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models
von: Huang, Yuhang, et al.
Veröffentlicht: (2024)
von: Huang, Yuhang, et al.
Veröffentlicht: (2024)
HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
von: Zeng, YangChen
Veröffentlicht: (2025)
von: Zeng, YangChen
Veröffentlicht: (2025)
Dual Mutual Learning Network with Global-local Awareness for RGB-D Salient Object Detection
von: Yi, Kang, et al.
Veröffentlicht: (2025)
von: Yi, Kang, et al.
Veröffentlicht: (2025)
Fine-grained Image Retrieval via Dual-Vision Adaptation
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models
von: Zhu, Hongyi, et al.
Veröffentlicht: (2024)
von: Zhu, Hongyi, et al.
Veröffentlicht: (2024)
Selective Vision-Language Subspace Projection for Few-shot CLIP
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks
von: Xu, Jingning, et al.
Veröffentlicht: (2026)
von: Xu, Jingning, et al.
Veröffentlicht: (2026)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
VideoSTF: Stress-Testing Output Repetition in Video Large Language Models
von: Cao, Yuxin, et al.
Veröffentlicht: (2026)
von: Cao, Yuxin, et al.
Veröffentlicht: (2026)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
von: Song, Zijie, et al.
Veröffentlicht: (2023)
von: Song, Zijie, et al.
Veröffentlicht: (2023)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
von: Zhou, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhou, Yuchen, et al.
Veröffentlicht: (2025)
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
von: Zhu, Haodong, et al.
Veröffentlicht: (2025)
von: Zhu, Haodong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025) -
Mitigating Image Captioning Hallucinations in Vision-Language Models
von: Zhao, Fei, et al.
Veröffentlicht: (2025) -
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
von: Wang, Xintong, et al.
Veröffentlicht: (2024) -
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
von: Fan, Linfeng, et al.
Veröffentlicht: (2026) -
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
von: Chen, Yangneng, et al.
Veröffentlicht: (2026)