From Retinal Evidence to Safe Decisions: RETINA-SAFE and ECRT for Hallucination Risk Triage in Medical LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Zhe, Xing, Wenpeng, Han, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LatentAudit: Real-Time White-Box Faithfulness Monitoring for Retrieval-Augmented Generation with Verifiable Deployment
by: Yu, Zhe, et al.
Published: (2026)
by: Yu, Zhe, et al.
Published: (2026)
TriLens: Per-Layer Logit-Lens Entropy for White-Box Hallucination Detection
by: Yang, Bohan, et al.
Published: (2026)
by: Yang, Bohan, et al.
Published: (2026)
Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning
by: Yu, Zhe, et al.
Published: (2026)
by: Yu, Zhe, et al.
Published: (2026)
The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context
by: Yu, Zhe, et al.
Published: (2026)
by: Yu, Zhe, et al.
Published: (2026)
SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
HGMF: A Hierarchical Gaussian Mixture Framework for Scalable Tool Invocation within the Model Context Protocol
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
Explainable AML Triage with LLMs: Evidence Retrieval and Counterfactual Checks
by: Torres, Dorothy, et al.
Published: (2026)
by: Torres, Dorothy, et al.
Published: (2026)
Language Models are Alignable Decision-Makers: Dataset and Application to the Medical Triage Domain
by: Hu, Brian, et al.
Published: (2024)
by: Hu, Brian, et al.
Published: (2024)
Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control
by: Yu, Zhe, et al.
Published: (2026)
by: Yu, Zhe, et al.
Published: (2026)
Silencing the Guardrails: Inference-Time Jailbreaking via Dynamic Contextual Representation Ablation
by: Xing, Wenpeng, et al.
Published: (2026)
by: Xing, Wenpeng, et al.
Published: (2026)
Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
by: Zhou, Yi, et al.
Published: (2025)
by: Zhou, Yi, et al.
Published: (2025)
Collaborative Medical Triage under Uncertainty: A Multi-Agent Dynamic Matching Approach
by: Cheng, Hongyan, et al.
Published: (2025)
by: Cheng, Hongyan, et al.
Published: (2025)
Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
by: Asadi, Mohammad, et al.
Published: (2026)
by: Asadi, Mohammad, et al.
Published: (2026)
Support Systems of Clinical Decisions in the Triage of the Emergency Department Using Artificial Intelligence: The Efficiency to Support Triage
by: Eleni Karlafti
Published: (2023)
by: Eleni Karlafti
Published: (2023)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
by: Lin, Shi, et al.
Published: (2024)
by: Lin, Shi, et al.
Published: (2024)
Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice
by: Doshi, Savan
Published: (2026)
by: Doshi, Savan
Published: (2026)
Leveraging Machine Learning Models to Predict the Outcome of Digital Medical Triage Interviews
by: Krylova, Sofia, et al.
Published: (2025)
by: Krylova, Sofia, et al.
Published: (2025)
Biased-Attention Guided Risk Prediction for Safe Decision-Making at Unsignalized Intersections
by: Dong, Chengyang, et al.
Published: (2025)
by: Dong, Chengyang, et al.
Published: (2025)
Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
by: Son, Yejin, et al.
Published: (2025)
by: Son, Yejin, et al.
Published: (2025)
Development of a Large Language Model-based Multi-Agent Clinical Decision Support System for Korean Triage and Acuity Scale (KTAS)-Based Triage and Treatment Planning in Emergency Departments
by: Han, Seungjun, et al.
Published: (2024)
by: Han, Seungjun, et al.
Published: (2024)
MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
by: Zhang, Jingxuan, et al.
Published: (2025)
by: Zhang, Jingxuan, et al.
Published: (2025)
Fairness in Healthcare Processes: A Quantitative Analysis of Decision Making in Triage
by: Andreswari, Rachmadita, et al.
Published: (2026)
by: Andreswari, Rachmadita, et al.
Published: (2026)
Do LLMs Triage Like Clinicians? A Dynamic Study of Outpatient Referral
by: Liu, Xiaoxiao, et al.
Published: (2025)
by: Liu, Xiaoxiao, et al.
Published: (2025)
Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs
by: Vaddi, Snehit, et al.
Published: (2026)
by: Vaddi, Snehit, et al.
Published: (2026)
Decision-Aware Trust Signal Alignment for SOC Alert Triage
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
by: Tan, Yixin, et al.
Published: (2025)
by: Tan, Yixin, et al.
Published: (2025)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
by: Li, Jin, et al.
Published: (2025)
by: Li, Jin, et al.
Published: (2025)
DIAP: A Decentralized Agent Identity Protocol with Zero-Knowledge Proofs and a Hybrid P2P Stack
by: Liu, Yuanjie, et al.
Published: (2025)
by: Liu, Yuanjie, et al.
Published: (2025)
Disentangling Deception and Hallucination Failures in LLMs
by: Lu, Haolang, et al.
Published: (2026)
by: Lu, Haolang, et al.
Published: (2026)
NLI:Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs Inference
by: Yu, Jiangyong, et al.
Published: (2026)
by: Yu, Jiangyong, et al.
Published: (2026)
ForgetMark: Stealthy Fingerprint Embedding via Targeted Unlearning in Language Models
by: Xu, Zhenhua, et al.
Published: (2026)
by: Xu, Zhenhua, et al.
Published: (2026)
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
by: Wang, Zhebo, et al.
Published: (2026)
by: Wang, Zhebo, et al.
Published: (2026)
Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering
by: Liu, Yaokun, et al.
Published: (2026)
by: Liu, Yaokun, et al.
Published: (2026)
SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents
by: Zhu, Kunlun, et al.
Published: (2025)
by: Zhu, Kunlun, et al.
Published: (2025)
SAFE-KD: Risk-Controlled Early-Exit Distillation for Vision Backbones
by: Khazem, Salim
Published: (2026)
by: Khazem, Salim
Published: (2026)
Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
by: Xu, Wenpeng
Published: (2026)
by: Xu, Wenpeng
Published: (2026)
Thinking, Faithful and Stable: Mitigating Hallucinations in LLMs
by: Zou, Chelsea, et al.
Published: (2025)
by: Zou, Chelsea, et al.
Published: (2025)
Implicit Geographic Inference in LLM Medical Triage: Language-Driven Disparities in Emergency Recommendations
by: Wong, Qi Han
Published: (2026)
by: Wong, Qi Han
Published: (2026)
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
by: Wang, Benlu, et al.
Published: (2025)
by: Wang, Benlu, et al.
Published: (2025)
Similar Items
-
LatentAudit: Real-Time White-Box Faithfulness Monitoring for Retrieval-Augmented Generation with Verifiable Deployment
by: Yu, Zhe, et al.
Published: (2026) -
TriLens: Per-Layer Logit-Lens Entropy for White-Box Hallucination Detection
by: Yang, Bohan, et al.
Published: (2026) -
Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning
by: Yu, Zhe, et al.
Published: (2026) -
The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context
by: Yu, Zhe, et al.
Published: (2026) -
SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
by: Xing, Wenpeng, et al.
Published: (2025)