Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Fu, Yuhan, Xie, Ruobing, Sun, Xingwu, Kang, Zhanhui, Li, Xirong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions
por: Fu, Yuhan, et al.
Publicado: (2024)
por: Fu, Yuhan, et al.
Publicado: (2024)
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
por: Compagnoni, Alberto, et al.
Publicado: (2025)
por: Compagnoni, Alberto, et al.
Publicado: (2025)
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
por: Fazli, Mehrdad, et al.
Publicado: (2025)
por: Fazli, Mehrdad, et al.
Publicado: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
por: Wang, Xintong, et al.
Publicado: (2024)
por: Wang, Xintong, et al.
Publicado: (2024)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
por: Deng, Ailin, et al.
Publicado: (2024)
por: Deng, Ailin, et al.
Publicado: (2024)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
por: Poppi, Tobia, et al.
Publicado: (2026)
por: Poppi, Tobia, et al.
Publicado: (2026)
Mitigating Image Captioning Hallucinations in Vision-Language Models
por: Zhao, Fei, et al.
Publicado: (2025)
por: Zhao, Fei, et al.
Publicado: (2025)
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
por: Liu, Fuxiao, et al.
Publicado: (2023)
por: Liu, Fuxiao, et al.
Publicado: (2023)
On the Audio Hallucinations in Large Audio-Video Language Models
por: Nishimura, Taichi, et al.
Publicado: (2024)
por: Nishimura, Taichi, et al.
Publicado: (2024)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
por: Fan, Linfeng, et al.
Publicado: (2026)
por: Fan, Linfeng, et al.
Publicado: (2026)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
por: Song, Jiale, et al.
Publicado: (2026)
por: Song, Jiale, et al.
Publicado: (2026)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
por: Deng, Jingyuan, et al.
Publicado: (2025)
por: Deng, Jingyuan, et al.
Publicado: (2025)
MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
por: Wang, Chenxi, et al.
Publicado: (2024)
por: Wang, Chenxi, et al.
Publicado: (2024)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
por: Wang, Xinran, et al.
Publicado: (2026)
por: Wang, Xinran, et al.
Publicado: (2026)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
por: Wu, Jiulong, et al.
Publicado: (2025)
por: Wu, Jiulong, et al.
Publicado: (2025)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
por: Zhang, Yudong, et al.
Publicado: (2024)
por: Zhang, Yudong, et al.
Publicado: (2024)
WordArt Designer API: User-Driven Artistic Typography Synthesis with Large Language Models on ModelScope
por: He, Jun-Yan, et al.
Publicado: (2024)
por: He, Jun-Yan, et al.
Publicado: (2024)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
por: Ding, Peng, et al.
Publicado: (2024)
por: Ding, Peng, et al.
Publicado: (2024)
Beyond Coarse-Grained Matching in Video-Text Retrieval
por: Chen, Aozhu, et al.
Publicado: (2024)
por: Chen, Aozhu, et al.
Publicado: (2024)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
por: Wu, Jiaying, et al.
Publicado: (2025)
por: Wu, Jiaying, et al.
Publicado: (2025)
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
por: Wu, Qiong, et al.
Publicado: (2024)
por: Wu, Qiong, et al.
Publicado: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
por: Ji, Yatai, et al.
Publicado: (2024)
por: Ji, Yatai, et al.
Publicado: (2024)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
por: Jiang, Chaoya, et al.
Publicado: (2024)
por: Jiang, Chaoya, et al.
Publicado: (2024)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
por: Bin, Yi, et al.
Publicado: (2024)
por: Bin, Yi, et al.
Publicado: (2024)
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
por: Chen, Yangneng, et al.
Publicado: (2026)
por: Chen, Yangneng, et al.
Publicado: (2026)
The Revolution of Multimodal Large Language Models: A Survey
por: Caffagni, Davide, et al.
Publicado: (2024)
por: Caffagni, Davide, et al.
Publicado: (2024)
Towards Training-free Multimodal Hate Localisation with Large Language Models
por: Sun, Yueming, et al.
Publicado: (2026)
por: Sun, Yueming, et al.
Publicado: (2026)
Mitigating Object Hallucination via Robust Local Perception Search
por: Gao, Zixian, et al.
Publicado: (2025)
por: Gao, Zixian, et al.
Publicado: (2025)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
por: Cheng, Zebang, et al.
Publicado: (2024)
por: Cheng, Zebang, et al.
Publicado: (2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
por: Liu, Jiazhen, et al.
Publicado: (2024)
por: Liu, Jiazhen, et al.
Publicado: (2024)
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
por: Bucciarelli, Davide, et al.
Publicado: (2024)
por: Bucciarelli, Davide, et al.
Publicado: (2024)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
por: Lu, Yifan, et al.
Publicado: (2025)
por: Lu, Yifan, et al.
Publicado: (2025)
Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
por: Zhao, Zhiyuan, et al.
Publicado: (2023)
por: Zhao, Zhiyuan, et al.
Publicado: (2023)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
por: Song, Shezheng, et al.
Publicado: (2023)
por: Song, Shezheng, et al.
Publicado: (2023)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
por: Zhao, Zhixian, et al.
Publicado: (2026)
por: Zhao, Zhixian, et al.
Publicado: (2026)
ChronusOmni: Improving Time Awareness of Omni Large Language Models
por: Chen, Yijing, et al.
Publicado: (2025)
por: Chen, Yijing, et al.
Publicado: (2025)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
por: Li, Jinyuan, et al.
Publicado: (2024)
por: Li, Jinyuan, et al.
Publicado: (2024)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
por: Zhang, Dongxu, et al.
Publicado: (2026)
por: Zhang, Dongxu, et al.
Publicado: (2026)
RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models
por: Mattioli, Gabriele, et al.
Publicado: (2026)
por: Mattioli, Gabriele, et al.
Publicado: (2026)
Ejemplares similares
-
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions
por: Fu, Yuhan, et al.
Publicado: (2024) -
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
por: Compagnoni, Alberto, et al.
Publicado: (2025) -
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
por: Fazli, Mehrdad, et al.
Publicado: (2025) -
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
por: Wang, Xintong, et al.
Publicado: (2024) -
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
por: Deng, Ailin, et al.
Publicado: (2024)