Skip \n: A Simple Method to Reduce Hallucination in Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Zongbo, Bai, Zechen, Mei, Haiyang, Xu, Qianli, Zhang, Changqing, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DOTA: Distributional Test-Time Adaptation of Vision-Language Models
von: Han, Zongbo, et al.
Veröffentlicht: (2024)
von: Han, Zongbo, et al.
Veröffentlicht: (2024)
Hallucination of Multimodal Large Language Models: A Survey
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
LOVA3: Learning to Visual Question Answering, Asking and Assessment
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)
Impossible Videos
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
von: Ma, Ruiqi, et al.
Veröffentlicht: (2025)
von: Ma, Ruiqi, et al.
Veröffentlicht: (2025)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
von: Wang, Weihang, et al.
Veröffentlicht: (2025)
von: Wang, Weihang, et al.
Veröffentlicht: (2025)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
von: Chang, Yue, et al.
Veröffentlicht: (2024)
von: Chang, Yue, et al.
Veröffentlicht: (2024)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
von: Fang, Hao, et al.
Veröffentlicht: (2025)
von: Fang, Hao, et al.
Veröffentlicht: (2025)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
von: Oh, Hongseok, et al.
Veröffentlicht: (2025)
von: Oh, Hongseok, et al.
Veröffentlicht: (2025)
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
von: Shang, Yuying, et al.
Veröffentlicht: (2024)
von: Shang, Yuying, et al.
Veröffentlicht: (2024)
A Survey on Hallucination in Large Vision-Language Models
von: Liu, Hanchao, et al.
Veröffentlicht: (2024)
von: Liu, Hanchao, et al.
Veröffentlicht: (2024)
Mitigating Multilingual Hallucination in Large Vision-Language Models
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
Benchmarking Deflection and Hallucination in Large Vision-Language Models
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2026)
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2026)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
von: Zhong, Weihong, et al.
Veröffentlicht: (2024)
von: Zhong, Weihong, et al.
Veröffentlicht: (2024)
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
von: Hartman, Max, et al.
Veröffentlicht: (2025)
von: Hartman, Max, et al.
Veröffentlicht: (2025)
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025)
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
von: Ha, Jiwoo, et al.
Veröffentlicht: (2026)
von: Ha, Jiwoo, et al.
Veröffentlicht: (2026)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025)
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025)
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
von: Guan, Tianrui, et al.
Veröffentlicht: (2023)
von: Guan, Tianrui, et al.
Veröffentlicht: (2023)
Multi-Object Hallucination in Vision-Language Models
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
von: Xu, Qinwu
Veröffentlicht: (2026)
von: Xu, Qinwu
Veröffentlicht: (2026)
Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models
von: Sanogo, Kassoum, et al.
Veröffentlicht: (2025)
von: Sanogo, Kassoum, et al.
Veröffentlicht: (2025)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DOTA: Distributional Test-Time Adaptation of Vision-Language Models
von: Han, Zongbo, et al.
Veröffentlicht: (2024) -
Hallucination of Multimodal Large Language Models: A Survey
von: Bai, Zechen, et al.
Veröffentlicht: (2024) -
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
von: Bai, Zechen, et al.
Veröffentlicht: (2025) -
LOVA3: Learning to Visual Question Answering, Asking and Assessment
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024) -
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)