Visual hallucination detection in large vision-language models via evidential conflict
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Tao, Liu, Zhekun, Wang, Rui, Zhang, Yang, Jing, Liping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Strong Transferable Adversarial Attacks via Ensembled Asymptotically Normal Distribution Learning
by: Fang, Zhengwei, et al.
Published: (2022)
by: Fang, Zhengwei, et al.
Published: (2022)
Visual representations in the human brain are aligned with large language models
by: Doerig, Adrien, et al.
Published: (2022)
by: Doerig, Adrien, et al.
Published: (2022)
evclust: Python library for evidential clustering
by: Soubeiga, Armel, et al.
Published: (2025)
by: Soubeiga, Armel, et al.
Published: (2025)
Triggering hallucinations in model-based MRI reconstruction via adversarial perturbations
by: Buğday, Suna, et al.
Published: (2026)
by: Buğday, Suna, et al.
Published: (2026)
Amortizing intractable inference in diffusion models for vision, language, and control
by: Venkatraman, Siddarth, et al.
Published: (2024)
by: Venkatraman, Siddarth, et al.
Published: (2024)
Predicting Distance matrix with large language models
by: Yang, Jiaxing
Published: (2024)
by: Yang, Jiaxing
Published: (2024)
ViSTa Dataset: Do vision-language models understand sequential tasks?
by: Wybitul, Evžen, et al.
Published: (2024)
by: Wybitul, Evžen, et al.
Published: (2024)
FlySearch: Exploring how vision-language models explore
by: Pardyl, Adam, et al.
Published: (2025)
by: Pardyl, Adam, et al.
Published: (2025)
Advancing vision-language models in front-end development via data synthesis
by: Ge, Tong, et al.
Published: (2025)
by: Ge, Tong, et al.
Published: (2025)
Reasoning emerges from constrained inference manifolds in large language models
by: Ma, Yanbiao, et al.
Published: (2026)
by: Ma, Yanbiao, et al.
Published: (2026)
What do vision-language models see in the context? Investigating multimodal in-context learning
by: Santos, Gabriel O. dos, et al.
Published: (2025)
by: Santos, Gabriel O. dos, et al.
Published: (2025)
Contrastive vision-language learning with paraphrasing and negation
by: Ngan, Kwun Ho, et al.
Published: (2025)
by: Ngan, Kwun Ho, et al.
Published: (2025)
A robust and scalable framework for hallucination detection in virtual tissue staining and digital pathology
by: Huang, Luzhe, et al.
Published: (2024)
by: Huang, Luzhe, et al.
Published: (2024)
BRAVE: Broadening the visual encoding of vision-language models
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
Bridging visual saliency and large language models for explainable deep learning in medical imaging
by: Nguezet, Paul Valery, et al.
Published: (2026)
by: Nguezet, Paul Valery, et al.
Published: (2026)
sFRC for assessing hallucinations in medical image restoration
by: Kc, Prabhat, et al.
Published: (2026)
by: Kc, Prabhat, et al.
Published: (2026)
Video Annotator: A framework for efficiently building video classifiers using vision-language models and active learning
by: Ziai, Amir, et al.
Published: (2024)
by: Ziai, Amir, et al.
Published: (2024)
Robust image classification with multi-modal large language models
by: Villani, Francesco, et al.
Published: (2024)
by: Villani, Francesco, et al.
Published: (2024)
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
GP-VLS: A general-purpose vision language model for surgery
by: Schmidgall, Samuel, et al.
Published: (2024)
by: Schmidgall, Samuel, et al.
Published: (2024)
Visual concept ranking uncovers medical shortcuts used by large multimodal models
by: Janizek, Joseph D., et al.
Published: (2026)
by: Janizek, Joseph D., et al.
Published: (2026)
Explainable artificial intelligence (XAI): from inherent explainability to large language models
by: Mumuni, Fuseini, et al.
Published: (2025)
by: Mumuni, Fuseini, et al.
Published: (2025)
Multi-label out-of-distribution detection via evidential learning
by: Aguilar, Eduardo, et al.
Published: (2025)
by: Aguilar, Eduardo, et al.
Published: (2025)
Open Set Domain Adaptation with Vision-language models via Gradient-aware Separation
by: Chen, Haoyang
Published: (2025)
by: Chen, Haoyang
Published: (2025)
Learning rigid-body simulators over implicit shapes for large-scale scenes and vision
by: Rubanova, Yulia, et al.
Published: (2024)
by: Rubanova, Yulia, et al.
Published: (2024)
PathAlign: A vision-language model for whole slide images in histopathology
by: Ahmed, Faruk, et al.
Published: (2024)
by: Ahmed, Faruk, et al.
Published: (2024)
1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit
by: Gao, Chang, et al.
Published: (2024)
by: Gao, Chang, et al.
Published: (2024)
Adapting to Distribution Shift by Visual Domain Prompt Generation
by: Chi, Zhixiang, et al.
Published: (2024)
by: Chi, Zhixiang, et al.
Published: (2024)
Cross-modal linkage risk in clinical vision-language models
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
The Solution for the GAIIC2024 RGB-TIR object detection Challenge
by: Wu, Xiangyu, et al.
Published: (2024)
by: Wu, Xiangyu, et al.
Published: (2024)
COOD: Combined out-of-distribution detection using multiple measures for anomaly & novel class detection in large-scale hierarchical classification
by: Hogeweg, L. E., et al.
Published: (2024)
by: Hogeweg, L. E., et al.
Published: (2024)
Multimodal Machine Translation with Visual Scene Graph Pruning
by: Lu, Chenyu, et al.
Published: (2025)
by: Lu, Chenyu, et al.
Published: (2025)
Aligning Visual Contrastive learning models via Preference Optimization
by: Afzali, Amirabbas, et al.
Published: (2024)
by: Afzali, Amirabbas, et al.
Published: (2024)
Linking heterogeneous microstructure informatics with expert characterization knowledge through customized and hybrid vision-language representations for industrial qualification
by: Safdar, Mutahar, et al.
Published: (2025)
by: Safdar, Mutahar, et al.
Published: (2025)
Tarsier: Recipes for Training and Evaluating Large Video Description Models
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
Look, Remember and Reason: Grounded reasoning in videos with language models
by: Bhattacharyya, Apratim, et al.
Published: (2023)
by: Bhattacharyya, Apratim, et al.
Published: (2023)
Enhancing Diffusion-based Point Cloud Generation with Smoothness Constraint
by: Li, Yukun, et al.
Published: (2024)
by: Li, Yukun, et al.
Published: (2024)
ALoRE: Efficient Visual Adaptation via Aggregating Low Rank Experts
by: Du, Sinan, et al.
Published: (2024)
by: Du, Sinan, et al.
Published: (2024)
Image anomaly detection and prediction scheme based on SSA optimized ResNet50-BiGRU model
by: Wan, Qianhui, et al.
Published: (2024)
by: Wan, Qianhui, et al.
Published: (2024)
When Classes Evolve: A Benchmark and Framework for Stage-Aware Class-Incremental Learning
by: Zhang, Zheng, et al.
Published: (2026)
by: Zhang, Zheng, et al.
Published: (2026)
Similar Items
-
Strong Transferable Adversarial Attacks via Ensembled Asymptotically Normal Distribution Learning
by: Fang, Zhengwei, et al.
Published: (2022) -
Visual representations in the human brain are aligned with large language models
by: Doerig, Adrien, et al.
Published: (2022) -
evclust: Python library for evidential clustering
by: Soubeiga, Armel, et al.
Published: (2025) -
Triggering hallucinations in model-based MRI reconstruction via adversarial perturbations
by: Buğday, Suna, et al.
Published: (2026) -
Amortizing intractable inference in diffusion models for vision, language, and control
by: Venkatraman, Siddarth, et al.
Published: (2024)