Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Back, Kyungryul, Park, Seongbeom, Kim, Milim, Kwon, Mincheol, Lee, SangHyeok, Lee, Hyunyoung, Cho, Junhee, Park, Seunghyun, Kim, Jinkyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915557444419584
author Back, Kyungryul
Park, Seongbeom
Kim, Milim
Kwon, Mincheol
Lee, SangHyeok
Lee, Hyunyoung
Cho, Junhee
Park, Seunghyun
Kim, Jinkyu
author_facet Back, Kyungryul
Park, Seongbeom
Kim, Milim
Kwon, Mincheol
Lee, SangHyeok
Lee, Hyunyoung
Cho, Junhee
Park, Seunghyun
Kim, Jinkyu
contents Large Vision-Language Models (LVLMs) have recently shown promising results on various multimodal tasks, even achieving human-comparable performance in certain cases. Nevertheless, LVLMs remain prone to hallucinations -- they often rely heavily on a single modality or memorize training data without properly grounding their outputs. To address this, we propose a training-free, tri-layer contrastive decoding with watermarking, which proceeds in three steps: (1) select a mature layer and an amateur layer among the decoding layers, (2) identify a pivot layer using a watermark-related question to assess whether the layer is visually well-grounded, and (3) apply tri-layer contrastive decoding to generate the final output. Experiments on public benchmarks such as POPE, MME and AMBER demonstrate that our method achieves state-of-the-art performance in reducing hallucinations in LVLMs and generates more visually grounded responses.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
Back, Kyungryul
Park, Seongbeom
Kim, Milim
Kwon, Mincheol
Lee, SangHyeok
Lee, Hyunyoung
Cho, Junhee
Park, Seunghyun
Kim, Jinkyu
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Vision-Language Models (LVLMs) have recently shown promising results on various multimodal tasks, even achieving human-comparable performance in certain cases. Nevertheless, LVLMs remain prone to hallucinations -- they often rely heavily on a single modality or memorize training data without properly grounding their outputs. To address this, we propose a training-free, tri-layer contrastive decoding with watermarking, which proceeds in three steps: (1) select a mature layer and an amateur layer among the decoding layers, (2) identify a pivot layer using a watermark-related question to assess whether the layer is visually well-grounded, and (3) apply tri-layer contrastive decoding to generate the final output. Experiments on public benchmarks such as POPE, MME and AMBER demonstrate that our method achieves state-of-the-art performance in reducing hallucinations in LVLMs and generates more visually grounded responses.
title Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.14304