SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhaoxu, Kong, Chenqi, Yu, Yi, Wu, Qiangqiang, Jiang, Xinghao, Cheung, Ngai-Man, Wen, Bihan, Kot, Alex, Jiang, Xudong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding
by: Li, Zhaoxu, et al.
Published: (2026)
by: Li, Zhaoxu, et al.
Published: (2026)
On the Adversarial Robustness of 3D Large Vision-Language Models
by: Liu, Chao, et al.
Published: (2026)
by: Liu, Chao, et al.
Published: (2026)
Adversarial Prompt Injection Attack on Multimodal Large Language Models
by: Ding, Meiwen, et al.
Published: (2026)
by: Ding, Meiwen, et al.
Published: (2026)
Sparse by Rule: Probability-Based N:M Pruning for Spiking Neural Networks
by: Ye, Shuhan, et al.
Published: (2025)
by: Ye, Shuhan, et al.
Published: (2025)
Learning from Dense Events: Towards Fast Spiking Neural Networks Training via Event Dataset Distillation
by: Ye, Shuhan, et al.
Published: (2025)
by: Ye, Shuhan, et al.
Published: (2025)
Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI
by: Nguyen, Phat, et al.
Published: (2025)
by: Nguyen, Phat, et al.
Published: (2025)
SegPoint: Segment Any Point Cloud via Large Language Model
by: He, Shuting, et al.
Published: (2024)
by: He, Shuting, et al.
Published: (2024)
Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking
by: Wu, Qiangqiang, et al.
Published: (2025)
by: Wu, Qiangqiang, et al.
Published: (2025)
From Chaos to Clarity: 3DGS in the Dark
by: Li, Zhihao, et al.
Published: (2024)
by: Li, Zhihao, et al.
Published: (2024)
SAVER: Selective As-Needed Vision Evidence for Multimodal Information Extraction
by: Hu, Miaobo, et al.
Published: (2026)
by: Hu, Miaobo, et al.
Published: (2026)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
by: Lee, Yi-Lun, et al.
Published: (2024)
by: Lee, Yi-Lun, et al.
Published: (2024)
S-Adapter: Generalizing Vision Transformer for Face Anti-Spoofing with Statistical Tokens
by: Cai, Rizhao, et al.
Published: (2023)
by: Cai, Rizhao, et al.
Published: (2023)
Open-Set Deepfake Detection: A Parameter-Efficient Adaptation Method with Forgery Style Mixture
by: Kong, Chenqi, et al.
Published: (2024)
by: Kong, Chenqi, et al.
Published: (2024)
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
by: Chen, Xinrong, et al.
Published: (2026)
by: Chen, Xinrong, et al.
Published: (2026)
Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
by: Zhang, Keyang, et al.
Published: (2025)
by: Zhang, Keyang, et al.
Published: (2025)
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024)
by: Wijaya, Robert, et al.
Published: (2024)
Feature-Space Smoothing: Certified Robustness of Deep Representations
by: Xia, Song, et al.
Published: (2026)
by: Xia, Song, et al.
Published: (2026)
Generalized Face Forgery Detection via Adaptive Learning for Pre-trained Vision Transformer
by: Luo, Anwei, et al.
Published: (2023)
by: Luo, Anwei, et al.
Published: (2023)
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
by: Zhang, Chengsheng, et al.
Published: (2026)
by: Zhang, Chengsheng, et al.
Published: (2026)
Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination
by: Yang, Dingchen, et al.
Published: (2024)
by: Yang, Dingchen, et al.
Published: (2024)
Frequency Masking for Universal Deepfake Detection
by: Doloriel, Chandler Timm, et al.
Published: (2024)
by: Doloriel, Chandler Timm, et al.
Published: (2024)
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
by: Jiang, Nick, et al.
Published: (2024)
by: Jiang, Nick, et al.
Published: (2024)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
by: Chen, Boqi, et al.
Published: (2026)
by: Chen, Boqi, et al.
Published: (2026)
When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
by: Lu, Hui, et al.
Published: (2025)
by: Lu, Hui, et al.
Published: (2025)
Mitigating Object Hallucinations via Sentence-Level Early Intervention
by: Peng, Shangpin, et al.
Published: (2025)
by: Peng, Shangpin, et al.
Published: (2025)
DP-IQA: Utilizing Diffusion Prior for Blind Image Quality Assessment in the Wild
by: Fu, Honghao, et al.
Published: (2024)
by: Fu, Honghao, et al.
Published: (2024)
Random Erasing vs. Model Inversion: A Promising Defense or a False Hope?
by: Tran, Viet-Hung, et al.
Published: (2024)
by: Tran, Viet-Hung, et al.
Published: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
ForensicsSAM: Toward Robust and Unified Image Forgery Detection and Localization Resisting to Adversarial Attack
by: Peng, Rongxuan, et al.
Published: (2025)
by: Peng, Rongxuan, et al.
Published: (2025)
Pixel-Inconsistency Modeling for Image Manipulation Localization
by: Kong, Chenqi, et al.
Published: (2023)
by: Kong, Chenqi, et al.
Published: (2023)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
by: Li, Qiming, et al.
Published: (2026)
by: Li, Qiming, et al.
Published: (2026)
AIR: Zero-shot Generative Model Adaptation with Iterative Refinement
by: Liu, Guimeng, et al.
Published: (2025)
by: Liu, Guimeng, et al.
Published: (2025)
Text to Image Generation and Editing: A Survey
by: Yang, Pengfei, et al.
Published: (2025)
by: Yang, Pengfei, et al.
Published: (2025)
Urban Air Temperature Prediction using Conditional Diffusion Models
by: Dai, Siyang, et al.
Published: (2024)
by: Dai, Siyang, et al.
Published: (2024)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
by: Lu, Yifan, et al.
Published: (2025)
by: Lu, Yifan, et al.
Published: (2025)
Vision Transformer Neural Architecture Search for Out-of-Distribution Generalization: Benchmark and Insights
by: Ho, Sy-Tuyen, et al.
Published: (2025)
by: Ho, Sy-Tuyen, et al.
Published: (2025)
VORD: Visual Ordinal Calibration for Mitigating Object Hallucinations in Large Vision-Language Models
by: Neo, Dexter, et al.
Published: (2024)
by: Neo, Dexter, et al.
Published: (2024)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
by: Lee, Jihoon, et al.
Published: (2025)
by: Lee, Jihoon, et al.
Published: (2025)
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
by: Yang, Chengxu, et al.
Published: (2026)
by: Yang, Chengxu, et al.
Published: (2026)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
by: Jiang, Yubo, et al.
Published: (2026)
by: Jiang, Yubo, et al.
Published: (2026)
Similar Items
-
SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding
by: Li, Zhaoxu, et al.
Published: (2026) -
On the Adversarial Robustness of 3D Large Vision-Language Models
by: Liu, Chao, et al.
Published: (2026) -
Adversarial Prompt Injection Attack on Multimodal Large Language Models
by: Ding, Meiwen, et al.
Published: (2026) -
Sparse by Rule: Probability-Based N:M Pruning for Spiking Neural Networks
by: Ye, Shuhan, et al.
Published: (2025) -
Learning from Dense Events: Towards Fast Spiking Neural Networks Training via Event Dataset Distillation
by: Ye, Shuhan, et al.
Published: (2025)