From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Shang, Yuying, Zeng, Xinyi, Zhu, Yutao, Yang, Xiao, Fang, Zhengwei, Zhang, Jingyuan, Chen, Jiawei, Liu, Zinan, Tian, Yu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
di: Zeng, Xinyi, et al.
Pubblicazione: (2024)
di: Zeng, Xinyi, et al.
Pubblicazione: (2024)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models
di: Zheng, Lifan, et al.
Pubblicazione: (2026)
di: Zheng, Lifan, et al.
Pubblicazione: (2026)
One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models
di: Zhu, Yutao, et al.
Pubblicazione: (2024)
di: Zhu, Yutao, et al.
Pubblicazione: (2024)
Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
di: Ma, Ruiqi, et al.
Pubblicazione: (2025)
di: Ma, Ruiqi, et al.
Pubblicazione: (2025)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
di: Zhou, Yiyang, et al.
Pubblicazione: (2023)
di: Zhou, Yiyang, et al.
Pubblicazione: (2023)
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
di: Arif, Kazi Hasan Ibn, et al.
Pubblicazione: (2025)
di: Arif, Kazi Hasan Ibn, et al.
Pubblicazione: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
di: An, Wenbin, et al.
Pubblicazione: (2024)
di: An, Wenbin, et al.
Pubblicazione: (2024)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
di: Geigle, Gregor, et al.
Pubblicazione: (2024)
di: Geigle, Gregor, et al.
Pubblicazione: (2024)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
di: Jing, Liqiang, et al.
Pubblicazione: (2025)
di: Jing, Liqiang, et al.
Pubblicazione: (2025)
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
di: Ren, Lingfeng, et al.
Pubblicazione: (2026)
di: Ren, Lingfeng, et al.
Pubblicazione: (2026)
LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models
di: Sun, Yizheng, et al.
Pubblicazione: (2025)
di: Sun, Yizheng, et al.
Pubblicazione: (2025)
Dynamic Token Reweighting for Robust Vision-Language Models
di: Jiang, Tanqiu, et al.
Pubblicazione: (2025)
di: Jiang, Tanqiu, et al.
Pubblicazione: (2025)
Multi-Object Hallucination in Vision-Language Models
di: Chen, Xuweiyi, et al.
Pubblicazione: (2024)
di: Chen, Xuweiyi, et al.
Pubblicazione: (2024)
Rethinking the Reliability of Multi-agent System: A Perspective from Byzantine Fault Tolerance
di: Zheng, Lifan, et al.
Pubblicazione: (2025)
di: Zheng, Lifan, et al.
Pubblicazione: (2025)
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
di: Chen, Junzhe, et al.
Pubblicazione: (2024)
di: Chen, Junzhe, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow
di: Bai, Jiaqi, et al.
Pubblicazione: (2025)
di: Bai, Jiaqi, et al.
Pubblicazione: (2025)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
di: Zhang, Ce, et al.
Pubblicazione: (2025)
di: Zhang, Ce, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
di: Lu, Yifan, et al.
Pubblicazione: (2025)
di: Lu, Yifan, et al.
Pubblicazione: (2025)
From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
di: Zhu, Mingcheng, et al.
Pubblicazione: (2026)
di: Zhu, Mingcheng, et al.
Pubblicazione: (2026)
Maintaining Informative Coherence: Migrating Hallucinations in Large Language Models via Absorbing Markov Chains
di: Wu, Jiemin, et al.
Pubblicazione: (2024)
di: Wu, Jiemin, et al.
Pubblicazione: (2024)
Revisiting Graph-Tokenizing Large Language Models: A Systematic Evaluation of Graph Token Understanding
di: Zhang, Zhongjian, et al.
Pubblicazione: (2026)
di: Zhang, Zhongjian, et al.
Pubblicazione: (2026)
Mitigating Multilingual Hallucination in Large Vision-Language Models
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
di: Lovenia, Holy, et al.
Pubblicazione: (2023)
di: Lovenia, Holy, et al.
Pubblicazione: (2023)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
di: Fang, Hao, et al.
Pubblicazione: (2025)
di: Fang, Hao, et al.
Pubblicazione: (2025)
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
di: Wu, Junfei, et al.
Pubblicazione: (2024)
di: Wu, Junfei, et al.
Pubblicazione: (2024)
Reference-free Hallucination Detection for Large Vision-Language Models
di: Li, Qing, et al.
Pubblicazione: (2024)
di: Li, Qing, et al.
Pubblicazione: (2024)
Can Large Language Models Understand Preferences in Personalized Recommendation?
di: Tan, Zhaoxuan, et al.
Pubblicazione: (2025)
di: Tan, Zhaoxuan, et al.
Pubblicazione: (2025)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
di: Woo, Sangmin, et al.
Pubblicazione: (2025)
di: Woo, Sangmin, et al.
Pubblicazione: (2025)
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025)
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025)
Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance
di: Zhao, Linxi, et al.
Pubblicazione: (2024)
di: Zhao, Linxi, et al.
Pubblicazione: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
di: Li, Bin, et al.
Pubblicazione: (2025)
di: Li, Bin, et al.
Pubblicazione: (2025)
TLDR: Token-Level Detective Reward Model for Large Vision Language Models
di: Fu, Deqing, et al.
Pubblicazione: (2024)
di: Fu, Deqing, et al.
Pubblicazione: (2024)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
di: Ye, Zekai, et al.
Pubblicazione: (2025)
di: Ye, Zekai, et al.
Pubblicazione: (2025)
First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
di: Ha, Jiwoo, et al.
Pubblicazione: (2026)
di: Ha, Jiwoo, et al.
Pubblicazione: (2026)
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
di: Xiao, Wenyi, et al.
Pubblicazione: (2024)
di: Xiao, Wenyi, et al.
Pubblicazione: (2024)
Truth-Aware Context Selection: Mitigating Hallucinations of Large Language Models Being Misled by Untruthful Contexts
di: Yu, Tian, et al.
Pubblicazione: (2024)
di: Yu, Tian, et al.
Pubblicazione: (2024)
Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
di: Li, Yanhong, et al.
Pubblicazione: (2025)
di: Li, Yanhong, et al.
Pubblicazione: (2025)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
di: Wang, Zhaowei, et al.
Pubblicazione: (2024)
di: Wang, Zhaowei, et al.
Pubblicazione: (2024)
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
di: Duan, Jinhao, et al.
Pubblicazione: (2025)
di: Duan, Jinhao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
di: Zeng, Xinyi, et al.
Pubblicazione: (2024) -
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
di: Seo, Hoigi, et al.
Pubblicazione: (2025) -
DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models
di: Zheng, Lifan, et al.
Pubblicazione: (2026) -
One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models
di: Zhu, Yutao, et al.
Pubblicazione: (2024) -
Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
di: Ma, Ruiqi, et al.
Pubblicazione: (2025)