Suppressing VLM Hallucinations with Spectral Representation Filtering
Fuente:
arXiv
Saved in:
| Main Authors: | Ali, Ameen, Zoabi, Tamim, Wolf, Lior |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Domain-Generalizable Multiple-Domain Clustering
by: Rozner, Amit, et al.
Published: (2023)
by: Rozner, Amit, et al.
Published: (2023)
Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression
by: Sarkar, Sreetama, et al.
Published: (2025)
by: Sarkar, Sreetama, et al.
Published: (2025)
A Novel CNN Gradient Boosting Ensemble for Guava Disease Detection
by: Rijon, Tamim Ahasan, et al.
Published: (2025)
by: Rijon, Tamim Ahasan, et al.
Published: (2025)
Understanding Dataset Distillation via Spectral Filtering
by: Bo, Deyu, et al.
Published: (2025)
by: Bo, Deyu, et al.
Published: (2025)
Speech-Synchronized Whiteboard Generation via VLM-Driven Structured Drawing Representations
by: Prasad, Suraj, et al.
Published: (2026)
by: Prasad, Suraj, et al.
Published: (2026)
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
by: Jiang, Nick, et al.
Published: (2024)
by: Jiang, Nick, et al.
Published: (2024)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
DocVLM: Make Your VLM an Efficient Reader
by: Nacson, Mor Shpigel, et al.
Published: (2024)
by: Nacson, Mor Shpigel, et al.
Published: (2024)
Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation
by: Banerjee, Oishi, et al.
Published: (2024)
by: Banerjee, Oishi, et al.
Published: (2024)
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
by: Wu, Zhenkai, et al.
Published: (2025)
by: Wu, Zhenkai, et al.
Published: (2025)
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities
by: Islam, Chashi Mahiul, et al.
Published: (2025)
by: Islam, Chashi Mahiul, et al.
Published: (2025)
Inductive Gradient Adjustment For Spectral Bias In Implicit Neural Representations
by: Shi, Kexuan, et al.
Published: (2024)
by: Shi, Kexuan, et al.
Published: (2024)
VIBE: Can a VLM Read the Room?
by: Chakraborty, Tania, et al.
Published: (2025)
by: Chakraborty, Tania, et al.
Published: (2025)
Large VLM-based Stylized Sports Captioning
by: Dhar, Sauptik, et al.
Published: (2025)
by: Dhar, Sauptik, et al.
Published: (2025)
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models
by: Zafar, Oz, et al.
Published: (2024)
by: Zafar, Oz, et al.
Published: (2024)
BMIP: Bi-directional Modality Interaction Prompt Learning for VLM
by: Lv, Song-Lin, et al.
Published: (2025)
by: Lv, Song-Lin, et al.
Published: (2025)
RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
by: Hu, Suhang, et al.
Published: (2025)
by: Hu, Suhang, et al.
Published: (2025)
BendVLM: Test-Time Debiasing of Vision-Language Embeddings
by: Gerych, Walter, et al.
Published: (2024)
by: Gerych, Walter, et al.
Published: (2024)
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
by: Li, Muyang, et al.
Published: (2026)
by: Li, Muyang, et al.
Published: (2026)
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
by: Liu, Jiajin, et al.
Published: (2026)
by: Liu, Jiajin, et al.
Published: (2026)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
by: Bao, Chen, et al.
Published: (2024)
by: Bao, Chen, et al.
Published: (2024)
ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
by: Li, Junxian, et al.
Published: (2024)
by: Li, Junxian, et al.
Published: (2024)
Progressive Alignment with VLM-LLM Feature to Augment Defect Classification for the ASE Dataset
by: Hsu, Chih-Chung, et al.
Published: (2024)
by: Hsu, Chih-Chung, et al.
Published: (2024)
Mitigating Diffusion Model Hallucinations with Dynamic Guidance
by: Triaridis, Kostas, et al.
Published: (2025)
by: Triaridis, Kostas, et al.
Published: (2025)
Suppressing Uncertainty in Gaze Estimation
by: Wang, Shijing, et al.
Published: (2024)
by: Wang, Shijing, et al.
Published: (2024)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
by: Sarwar, Nobin
Published: (2025)
by: Sarwar, Nobin
Published: (2025)
Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting
by: Zhong, Siru, et al.
Published: (2025)
by: Zhong, Siru, et al.
Published: (2025)
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
DiffFinger: Advancing Synthetic Fingerprint Generation through Denoising Diffusion Probabilistic Models
by: Grabovski, Freddie, et al.
Published: (2024)
by: Grabovski, Freddie, et al.
Published: (2024)
TerraMAE: Learning Spatial-Spectral Representations from Hyperspectral Earth Observation Data via Adaptive Masked Autoencoders
by: Faruk, Tanjim Bin, et al.
Published: (2025)
by: Faruk, Tanjim Bin, et al.
Published: (2025)
Coherence Awareness in Diffractive Neural Networks
by: Kleiner, Matan, et al.
Published: (2024)
by: Kleiner, Matan, et al.
Published: (2024)
Illumination Angular Spectrum Encoding for Controlling the Functionality of Diffractive Networks
by: Kleiner, Matan, et al.
Published: (2026)
by: Kleiner, Matan, et al.
Published: (2026)
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
by: Dong, Bowen, et al.
Published: (2025)
by: Dong, Bowen, et al.
Published: (2025)
Detecting and Preventing Hallucinations in Large Vision Language Models
by: Gunjal, Anisha, et al.
Published: (2023)
by: Gunjal, Anisha, et al.
Published: (2023)
Filter Like You Test: Data-Driven Data Filtering for CLIP Pretraining
by: Shechter, Mikey, et al.
Published: (2025)
by: Shechter, Mikey, et al.
Published: (2025)
FMCL: Class-Aware Client Clustering with Foundation Model Representations for Heterogeneous Federated Learning
by: Ali, Mahad, et al.
Published: (2026)
by: Ali, Mahad, et al.
Published: (2026)
SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
by: Yi, Huahui, et al.
Published: (2025)
by: Yi, Huahui, et al.
Published: (2025)
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
by: Sivakumar, Anushka, et al.
Published: (2025)
by: Sivakumar, Anushka, et al.
Published: (2025)
Similar Items
-
Domain-Generalizable Multiple-Domain Clustering
by: Rozner, Amit, et al.
Published: (2023) -
Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression
by: Sarkar, Sreetama, et al.
Published: (2025) -
A Novel CNN Gradient Boosting Ensemble for Guava Disease Detection
by: Rijon, Tamim Ahasan, et al.
Published: (2025) -
Understanding Dataset Distillation via Spectral Filtering
by: Bo, Deyu, et al.
Published: (2025) -
Speech-Synchronized Whiteboard Generation via VLM-Driven Structured Drawing Representations
by: Prasad, Suraj, et al.
Published: (2026)