SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Sangha, Yoo, Seungryong, Mok, Jisoo, Yoon, Sungroh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
von: Park, Sangha, et al.
Veröffentlicht: (2025)
von: Park, Sangha, et al.
Veröffentlicht: (2025)
RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
von: Oh, Yeongtak, et al.
Veröffentlicht: (2025)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2025)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
von: Lew, Jaihyun, et al.
Veröffentlicht: (2024)
von: Lew, Jaihyun, et al.
Veröffentlicht: (2024)
ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
von: Yin, Hao, et al.
Veröffentlicht: (2025)
von: Yin, Hao, et al.
Veröffentlicht: (2025)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
von: Hua, Zhenglin, et al.
Veröffentlicht: (2025)
von: Hua, Zhenglin, et al.
Veröffentlicht: (2025)
Contextualized Visual Personalization in Vision-Language Models
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation
von: Shin, Chaehun, et al.
Veröffentlicht: (2025)
von: Shin, Chaehun, et al.
Veröffentlicht: (2025)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification
von: Sun, Han, et al.
Veröffentlicht: (2026)
von: Sun, Han, et al.
Veröffentlicht: (2026)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
von: Park, Yeji, et al.
Veröffentlicht: (2024)
von: Park, Yeji, et al.
Veröffentlicht: (2024)
Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models
von: Sun, Li, et al.
Veröffentlicht: (2024)
von: Sun, Li, et al.
Veröffentlicht: (2024)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2025)
von: Woo, Sangmin, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions
von: Tan, Zhijie, et al.
Veröffentlicht: (2025)
von: Tan, Zhijie, et al.
Veröffentlicht: (2025)
Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations
von: Chen, Boxu, et al.
Veröffentlicht: (2025)
von: Chen, Boxu, et al.
Veröffentlicht: (2025)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
von: Liu, Yufang, et al.
Veröffentlicht: (2024)
von: Liu, Yufang, et al.
Veröffentlicht: (2024)
Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization
von: Song, Sanghyeob, et al.
Veröffentlicht: (2024)
von: Song, Sanghyeob, et al.
Veröffentlicht: (2024)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
von: Ha, Jiwoo, et al.
Veröffentlicht: (2026)
von: Ha, Jiwoo, et al.
Veröffentlicht: (2026)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
von: Lee, Jihoon, et al.
Veröffentlicht: (2025)
von: Lee, Jihoon, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
von: Zhu, Younan, et al.
Veröffentlicht: (2025)
von: Zhu, Younan, et al.
Veröffentlicht: (2025)
Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs
von: Yu, Liu, et al.
Veröffentlicht: (2025)
von: Yu, Liu, et al.
Veröffentlicht: (2025)
Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
von: Li, Yuanshuai, et al.
Veröffentlicht: (2025)
von: Li, Yuanshuai, et al.
Veröffentlicht: (2025)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
von: Jiang, Yubo, et al.
Veröffentlicht: (2026)
von: Jiang, Yubo, et al.
Veröffentlicht: (2026)
DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
von: Li, Shuo, et al.
Veröffentlicht: (2025)
von: Li, Shuo, et al.
Veröffentlicht: (2025)
GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations
von: Chen, Boyuan, et al.
Veröffentlicht: (2026)
von: Chen, Boyuan, et al.
Veröffentlicht: (2026)
Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation
von: Gong, Boyang, et al.
Veröffentlicht: (2026)
von: Gong, Boyang, et al.
Veröffentlicht: (2026)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
von: Zheng, Haojie, et al.
Veröffentlicht: (2024)
von: Zheng, Haojie, et al.
Veröffentlicht: (2024)
Countering the Over-Reliance Trap: Mitigating Object Hallucination for LVLMs via a Self-Validation Framework
von: Liu, Shiyu, et al.
Veröffentlicht: (2026)
von: Liu, Shiyu, et al.
Veröffentlicht: (2026)
Mitigating Open-Vocabulary Caption Hallucinations
von: Ben-Kish, Assaf, et al.
Veröffentlicht: (2023)
von: Ben-Kish, Assaf, et al.
Veröffentlicht: (2023)
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
von: Park, Seongheon, et al.
Veröffentlicht: (2025)
von: Park, Seongheon, et al.
Veröffentlicht: (2025)
Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy
von: Xie, Yutong, et al.
Veröffentlicht: (2026)
von: Xie, Yutong, et al.
Veröffentlicht: (2026)
URECA: Unique Region Caption Anything
von: Lim, Sangbeom, et al.
Veröffentlicht: (2025)
von: Lim, Sangbeom, et al.
Veröffentlicht: (2025)
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
von: Yin, Hao, et al.
Veröffentlicht: (2025)
von: Yin, Hao, et al.
Veröffentlicht: (2025)
Causal Interpretation of Sparse Autoencoder Features in Vision
von: Han, Sangyu, et al.
Veröffentlicht: (2025)
von: Han, Sangyu, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
von: Sun, Nan, et al.
Veröffentlicht: (2025)
von: Sun, Nan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
von: Lee, Saehyung, et al.
Veröffentlicht: (2024) -
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
von: Park, Sangha, et al.
Veröffentlicht: (2025) -
RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
von: Oh, Yeongtak, et al.
Veröffentlicht: (2025) -
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
von: Lew, Jaihyun, et al.
Veröffentlicht: (2024) -
ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
von: Yin, Hao, et al.
Veröffentlicht: (2025)