Consistency and Uncertainty: Identifying Unreliable Responses From Black-Box Vision-Language Models for Selective Visual Question Answering
Fuente:
arXiv
Salvato in:
| Autori principali: | Khan, Zaid, Fu, Yun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
di: Wieczorek, Tobias Jan, et al.
Pubblicazione: (2025)
di: Wieczorek, Tobias Jan, et al.
Pubblicazione: (2025)
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
di: Shah, Monika, et al.
Pubblicazione: (2025)
di: Shah, Monika, et al.
Pubblicazione: (2025)
Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
di: Hao, Dongze, et al.
Pubblicazione: (2024)
di: Hao, Dongze, et al.
Pubblicazione: (2024)
Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
di: Lan, Jian, et al.
Pubblicazione: (2025)
di: Lan, Jian, et al.
Pubblicazione: (2025)
Selectively Answering Visual Questions
di: Eisenschlos, Julian Martin, et al.
Pubblicazione: (2024)
di: Eisenschlos, Julian Martin, et al.
Pubblicazione: (2024)
Large Vision-Language Models for Remote Sensing Visual Question Answering
di: Siripong, Surasakdi, et al.
Pubblicazione: (2024)
di: Siripong, Surasakdi, et al.
Pubblicazione: (2024)
BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users
di: Cheng, Wanyin, et al.
Pubblicazione: (2025)
di: Cheng, Wanyin, et al.
Pubblicazione: (2025)
HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models
di: Bai, Xiangyu, et al.
Pubblicazione: (2026)
di: Bai, Xiangyu, et al.
Pubblicazione: (2026)
Fusion of Domain-Adapted Vision and Language Models for Medical Visual Question Answering
di: Ha, Cuong Nhat, et al.
Pubblicazione: (2024)
di: Ha, Cuong Nhat, et al.
Pubblicazione: (2024)
Adapting Lightweight Vision Language Models for Radiological Visual Question Answering
di: Shourya, Aditya, et al.
Pubblicazione: (2025)
di: Shourya, Aditya, et al.
Pubblicazione: (2025)
Few-Shot Image Classification and Segmentation as Visual Question Answering Using Vision-Language Models
di: Meng, Tian, et al.
Pubblicazione: (2024)
di: Meng, Tian, et al.
Pubblicazione: (2024)
SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving
di: Zheng, Peiru, et al.
Pubblicazione: (2024)
di: Zheng, Peiru, et al.
Pubblicazione: (2024)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
di: Woo, Sangmin, et al.
Pubblicazione: (2025)
di: Woo, Sangmin, et al.
Pubblicazione: (2025)
Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays
di: Cho, Yeongjae, et al.
Pubblicazione: (2024)
di: Cho, Yeongjae, et al.
Pubblicazione: (2024)
Test-Time Hinting for Black-Box Vision-Language Models
di: Hou, Kaihua, et al.
Pubblicazione: (2026)
di: Hou, Kaihua, et al.
Pubblicazione: (2026)
Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
di: Hartsock, Iryna, et al.
Pubblicazione: (2024)
di: Hartsock, Iryna, et al.
Pubblicazione: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
Object Retrieval for Visual Question Answering with Outside Knowledge
di: Kan, Shichao, et al.
Pubblicazione: (2024)
di: Kan, Shichao, et al.
Pubblicazione: (2024)
Research on Vision-Language Question Answering Models for Industrial Robots
di: Li, Ping, et al.
Pubblicazione: (2026)
di: Li, Ping, et al.
Pubblicazione: (2026)
Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering
di: Chen, Yixiong, et al.
Pubblicazione: (2025)
di: Chen, Yixiong, et al.
Pubblicazione: (2025)
3D Question Answering via only 2D Vision-Language Models
di: Wang, Fengyun, et al.
Pubblicazione: (2025)
di: Wang, Fengyun, et al.
Pubblicazione: (2025)
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering
di: Zhu, Yingjian, et al.
Pubblicazione: (2026)
di: Zhu, Yingjian, et al.
Pubblicazione: (2026)
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks
di: Lee, Jusung, et al.
Pubblicazione: (2024)
di: Lee, Jusung, et al.
Pubblicazione: (2024)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
di: Lagos, Maximiliano Hormazábal, et al.
Pubblicazione: (2025)
di: Lagos, Maximiliano Hormazábal, et al.
Pubblicazione: (2025)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
di: Zou, Bo, et al.
Pubblicazione: (2024)
di: Zou, Bo, et al.
Pubblicazione: (2024)
Object Attribute Matters in Visual Question Answering
di: Li, Peize, et al.
Pubblicazione: (2023)
di: Li, Peize, et al.
Pubblicazione: (2023)
Where do Large Vision-Language Models Look at when Answering Questions?
di: Xing, Xiaoying, et al.
Pubblicazione: (2025)
di: Xing, Xiaoying, et al.
Pubblicazione: (2025)
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
di: Wang, Yanling, et al.
Pubblicazione: (2025)
di: Wang, Yanling, et al.
Pubblicazione: (2025)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
di: Lim, Su Hyeon, et al.
Pubblicazione: (2024)
di: Lim, Su Hyeon, et al.
Pubblicazione: (2024)
Targeted Visual Prompting for Medical Visual Question Answering
di: Tascon-Morales, Sergio, et al.
Pubblicazione: (2024)
di: Tascon-Morales, Sergio, et al.
Pubblicazione: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
di: Ishmam, Md Farhan, et al.
Pubblicazione: (2024)
di: Ishmam, Md Farhan, et al.
Pubblicazione: (2024)
Visually Interpretable Subtask Reasoning for Visual Question Answering
di: Cheng, Yu, et al.
Pubblicazione: (2025)
di: Cheng, Yu, et al.
Pubblicazione: (2025)
A Two-Stage Multitask Vision-Language Framework for Explainable Crop Disease Visual Question Answering
di: Hossain, Md. Zahid, et al.
Pubblicazione: (2026)
di: Hossain, Md. Zahid, et al.
Pubblicazione: (2026)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
di: Kim, Hongyeob, et al.
Pubblicazione: (2025)
di: Kim, Hongyeob, et al.
Pubblicazione: (2025)
Questioning the Stability of Visual Question Answering
di: Rosenfeld, Amir, et al.
Pubblicazione: (2025)
di: Rosenfeld, Amir, et al.
Pubblicazione: (2025)
Trust the Unreliability: Inward Backward Dynamic Unreliability Driven Coreset Selection for Medical Image Classification
di: Liang, Yan, et al.
Pubblicazione: (2026)
di: Liang, Yan, et al.
Pubblicazione: (2026)
TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving
di: Hassani, Hossein, et al.
Pubblicazione: (2025)
di: Hassani, Hossein, et al.
Pubblicazione: (2025)
Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment
di: Chen, Jin, et al.
Pubblicazione: (2024)
di: Chen, Jin, et al.
Pubblicazione: (2024)
CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering
di: Zeng, Xiyin, et al.
Pubblicazione: (2026)
di: Zeng, Xiyin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
di: Wieczorek, Tobias Jan, et al.
Pubblicazione: (2025) -
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
di: Khan, Zaid, et al.
Pubblicazione: (2024) -
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
di: Shah, Monika, et al.
Pubblicazione: (2025) -
Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
di: Hao, Dongze, et al.
Pubblicazione: (2024) -
Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
di: Lan, Jian, et al.
Pubblicazione: (2025)