Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
Fuente:
arXiv
Guardado en:
| Autores principales: | Asadi, Mohammad, Nedaee, Tahoura, O'Sullivan, Jack W., Ashley, Euan, Adeli, Ehsan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MIRAGE: The Illusion of Visual Understanding
por: Asadi, Mohammad, et al.
Publicado: (2026)
por: Asadi, Mohammad, et al.
Publicado: (2026)
MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management
por: O'Sullivan, Jack W, et al.
Publicado: (2026)
por: O'Sullivan, Jack W, et al.
Publicado: (2026)
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
por: Khati, Dipin, et al.
Publicado: (2026)
por: Khati, Dipin, et al.
Publicado: (2026)
The First Token Knows: Single-Decode Confidence for Hallucination Detection
por: Gabriel, Mina
Publicado: (2026)
por: Gabriel, Mina
Publicado: (2026)
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
por: Zhou, Yue, et al.
Publicado: (2026)
por: Zhou, Yue, et al.
Publicado: (2026)
Refined Iterated Pareto Greedy for Energy-aware Hybrid Flowshop Scheduling with Blocking Constraints
por: Missaoui, Ahmed, et al.
Publicado: (2025)
por: Missaoui, Ahmed, et al.
Publicado: (2025)
The Initial Exploration Problem in Knowledge Graph Exploration
por: McNamara, Claire, et al.
Publicado: (2026)
por: McNamara, Claire, et al.
Publicado: (2026)
Lower Bound on Howard Policy Iteration for Deterministic Markov Decision Processes
por: Asadi, Ali, et al.
Publicado: (2025)
por: Asadi, Ali, et al.
Publicado: (2025)
AdaVid: Adaptive Video-Language Pretraining
por: Patel, Chaitanya, et al.
Publicado: (2025)
por: Patel, Chaitanya, et al.
Publicado: (2025)
Multimodal Climate Disinformation Detection: Integrating Vision-Language Models with External Knowledge Sources
por: Shamsabad, Marzieh Adeli, et al.
Publicado: (2026)
por: Shamsabad, Marzieh Adeli, et al.
Publicado: (2026)
TWIG: Towards pre-hoc Hyperparameter Optimisation and Cross-Graph Generalisation via Simulated KGE Models
por: Sardina, Jeffrey, et al.
Publicado: (2024)
por: Sardina, Jeffrey, et al.
Publicado: (2024)
Neuro-Symbolic Decoding of Neural Activity
por: Wang, Yanchen, et al.
Publicado: (2026)
por: Wang, Yanchen, et al.
Publicado: (2026)
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination
por: Liang, Qiyao, et al.
Publicado: (2026)
por: Liang, Qiyao, et al.
Publicado: (2026)
Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQA
por: Thakrar, Karishma, et al.
Publicado: (2025)
por: Thakrar, Karishma, et al.
Publicado: (2025)
KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA
por: Yang, Fan, et al.
Publicado: (2026)
por: Yang, Fan, et al.
Publicado: (2026)
Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis
por: Yeganeh, Yousef, et al.
Publicado: (2024)
por: Yeganeh, Yousef, et al.
Publicado: (2024)
Towards Democratization of Subspeciality Medical Expertise
por: O'Sullivan, Jack W., et al.
Publicado: (2024)
por: O'Sullivan, Jack W., et al.
Publicado: (2024)
Surrogate-Based Prevalence Measurement for Large-Scale A/B Testing
por: Xu, Zehao, et al.
Publicado: (2026)
por: Xu, Zehao, et al.
Publicado: (2026)
Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
por: Pandey, Ayush, et al.
Publicado: (2025)
por: Pandey, Ayush, et al.
Publicado: (2025)
From Retinal Evidence to Safe Decisions: RETINA-SAFE and ECRT for Hallucination Risk Triage in Medical LLMs
por: Yu, Zhe, et al.
Publicado: (2026)
por: Yu, Zhe, et al.
Publicado: (2026)
Beyond Retrieval: Modeling Confidence Decay and Deterministic Agentic Platforms in Generative Engine Optimization
por: Zhao, XinYu, et al.
Publicado: (2026)
por: Zhao, XinYu, et al.
Publicado: (2026)
Epistemic Filtering and Collective Hallucination: A Jury Theorem for Confidence-Calibrated Agents
por: Karge, Jonas
Publicado: (2026)
por: Karge, Jonas
Publicado: (2026)
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence
por: Wang, Xingyue, et al.
Publicado: (2026)
por: Wang, Xingyue, et al.
Publicado: (2026)
Towards Fast Algorithms for the Preference Consistency Problem Based on Hierarchical Models
por: George, Anne-Marie, et al.
Publicado: (2024)
por: George, Anne-Marie, et al.
Publicado: (2024)
Reasoning Transfer for an Extremely Low-Resource and Endangered Language: Bridging Languages Through Sample-Efficient Language Understanding
por: Tran, Khanh-Tung, et al.
Publicado: (2025)
por: Tran, Khanh-Tung, et al.
Publicado: (2025)
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
por: Tran, Khanh-Tung, et al.
Publicado: (2025)
por: Tran, Khanh-Tung, et al.
Publicado: (2025)
UCCIX: Irish-eXcellence Large Language Model
por: Tran, Khanh-Tung, et al.
Publicado: (2024)
por: Tran, Khanh-Tung, et al.
Publicado: (2024)
Optimising 4th-Order Runge-Kutta Methods: A Dynamic Heuristic Approach for Efficiency and Low Storage
por: Goodship, Gavin Lee, et al.
Publicado: (2025)
por: Goodship, Gavin Lee, et al.
Publicado: (2025)
Questionnaire meets LLM: A Benchmark and Empirical Study of Structural Skills for Understanding Questions and Responses
por: Nguyen, Duc-Hai, et al.
Publicado: (2025)
por: Nguyen, Duc-Hai, et al.
Publicado: (2025)
Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA
por: Yan, Qianqi, et al.
Publicado: (2024)
por: Yan, Qianqi, et al.
Publicado: (2024)
Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
por: Aminmansour, Farzane, et al.
Publicado: (2020)
por: Aminmansour, Farzane, et al.
Publicado: (2020)
GeoSAE: Geometric Prior-Guided Layer-Wise Sparse Autoencoder Annotation of Brain MRI Foundation Models
por: Nerrise, Favour, et al.
Publicado: (2026)
por: Nerrise, Favour, et al.
Publicado: (2026)
ToolVQA: A Dataset for Multi-step Reasoning VQA with External Tools
por: Yin, Shaofeng, et al.
Publicado: (2025)
por: Yin, Shaofeng, et al.
Publicado: (2025)
GAMMA-PD: Graph-based Analysis of Multi-Modal Motor Impairment Assessments in Parkinson's Disease
por: Nerrise, Favour, et al.
Publicado: (2024)
por: Nerrise, Favour, et al.
Publicado: (2024)
I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation
por: Zong, Haotian, et al.
Publicado: (2026)
por: Zong, Haotian, et al.
Publicado: (2026)
MedHal: An Evaluation Dataset for Medical Hallucination Detection
por: Mehenni, Gaya, et al.
Publicado: (2025)
por: Mehenni, Gaya, et al.
Publicado: (2025)
Qualitative Analysis of $ω$-Regular Objectives on Robust MDPs
por: Asadi, Ali, et al.
Publicado: (2025)
por: Asadi, Ali, et al.
Publicado: (2025)
MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book
por: Yip, Sau Lai, et al.
Publicado: (2025)
por: Yip, Sau Lai, et al.
Publicado: (2025)
Hallucination Detection and Hallucination Mitigation: An Investigation
por: Luo, Junliang, et al.
Publicado: (2024)
por: Luo, Junliang, et al.
Publicado: (2024)
Enhancing Hallucination Detection via Future Context
por: Lee, Joosung, et al.
Publicado: (2025)
por: Lee, Joosung, et al.
Publicado: (2025)
Ejemplares similares
-
MIRAGE: The Illusion of Visual Understanding
por: Asadi, Mohammad, et al.
Publicado: (2026) -
MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management
por: O'Sullivan, Jack W, et al.
Publicado: (2026) -
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
por: Khati, Dipin, et al.
Publicado: (2026) -
The First Token Knows: Single-Decode Confidence for Hallucination Detection
por: Gabriel, Mina
Publicado: (2026) -
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
por: Zhou, Yue, et al.
Publicado: (2026)