Counterfactual Cultural Cues Reduce Medical QA Accuracy in LLMs: Identifier vs Context Effects
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rezaei, Amirhossein Haji Mohammad, Shakeri, Zahra |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Public Health Text Annotation: Zero-Shot Learning vs. Crowdsourcing for Improved Efficiency and Labeling Accuracy
von: Kazari, Kamyar, et al.
Veröffentlicht: (2025)
von: Kazari, Kamyar, et al.
Veröffentlicht: (2025)
Detecting Bias and Enhancing Diagnostic Accuracy in Large Language Models for Healthcare
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2024)
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2024)
Can LLMs Explain Themselves Counterfactually?
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2025)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2025)
Aligning (Medical) LLMs for (Counterfactual) Fairness
von: Poulain, Raphael, et al.
Veröffentlicht: (2024)
von: Poulain, Raphael, et al.
Veröffentlicht: (2024)
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
von: Le, Chenqian, et al.
Veröffentlicht: (2025)
von: Le, Chenqian, et al.
Veröffentlicht: (2025)
Knowledge in Triples for LLMs: Enhancing Table QA Accuracy with Semantic Extraction
von: Sholehrasa, Hossein, et al.
Veröffentlicht: (2024)
von: Sholehrasa, Hossein, et al.
Veröffentlicht: (2024)
MuCo-KGC: Multi-Context-Aware Knowledge Graph Completion
von: Gul, Haji, et al.
Veröffentlicht: (2025)
von: Gul, Haji, et al.
Veröffentlicht: (2025)
Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
von: Ju, Li, et al.
Veröffentlicht: (2026)
von: Ju, Li, et al.
Veröffentlicht: (2026)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
von: Mo, Kaijie, et al.
Veröffentlicht: (2026)
von: Mo, Kaijie, et al.
Veröffentlicht: (2026)
Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
von: Yousefiramandi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Yousefiramandi, Amirhossein, et al.
Veröffentlicht: (2025)
Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical Knowledge
von: Rezaei, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Rezaei, Mohammad Reza, et al.
Veröffentlicht: (2025)
TyDi QA-WANA: A Benchmark for Information-Seeking Question Answering in Languages of West Asia and North Africa
von: Riley, Parker, et al.
Veröffentlicht: (2025)
von: Riley, Parker, et al.
Veröffentlicht: (2025)
MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis
von: Chen, Wenting, et al.
Veröffentlicht: (2026)
von: Chen, Wenting, et al.
Veröffentlicht: (2026)
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark
von: Wu, Jian, et al.
Veröffentlicht: (2024)
von: Wu, Jian, et al.
Veröffentlicht: (2024)
Assessing GPTZero's Accuracy in Identifying AI vs. Human-Written Essays
von: Dik, Selin, et al.
Veröffentlicht: (2025)
von: Dik, Selin, et al.
Veröffentlicht: (2025)
MuCoS: Efficient Drug-Target Prediction through Multi-Context-Aware Sampling
von: Gul, Haji, et al.
Veröffentlicht: (2025)
von: Gul, Haji, et al.
Veröffentlicht: (2025)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
von: Dineen, Jacob, et al.
Veröffentlicht: (2025)
von: Dineen, Jacob, et al.
Veröffentlicht: (2025)
Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA
von: Testoni, Alberto, et al.
Veröffentlicht: (2026)
von: Testoni, Alberto, et al.
Veröffentlicht: (2026)
Long Context vs. RAG for LLMs: An Evaluation and Revisits
von: Li, Xinze, et al.
Veröffentlicht: (2024)
von: Li, Xinze, et al.
Veröffentlicht: (2024)
ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities
von: Xu, Peng, et al.
Veröffentlicht: (2024)
von: Xu, Peng, et al.
Veröffentlicht: (2024)
Counterfactual Generation with Identifiability Guarantees
von: Yan, Hanqi, et al.
Veröffentlicht: (2024)
von: Yan, Hanqi, et al.
Veröffentlicht: (2024)
Vendi-RAG: Adaptively Trading-Off Diversity And Quality Significantly Improves Retrieval Augmented Generation With LLMs
von: Rezaei, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Rezaei, Mohammad Reza, et al.
Veröffentlicht: (2025)
How Well Do LLMs Identify Cultural Unity in Diversity?
von: Li, Jialin, et al.
Veröffentlicht: (2024)
von: Li, Jialin, et al.
Veröffentlicht: (2024)
Identifying Shopping Intent in Product QA for Proactive Recommendations
von: Fetahu, Besnik, et al.
Veröffentlicht: (2024)
von: Fetahu, Besnik, et al.
Veröffentlicht: (2024)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
ETT: Expanding the Long Context Understanding Capability of LLMs at Test-Time
von: Zahirnia, Kiarash, et al.
Veröffentlicht: (2025)
von: Zahirnia, Kiarash, et al.
Veröffentlicht: (2025)
MedConceptsQA: Open Source Medical Concepts QA Benchmark
von: Shoham, Ofir Ben, et al.
Veröffentlicht: (2024)
von: Shoham, Ofir Ben, et al.
Veröffentlicht: (2024)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
von: Chen, Jingyi, et al.
Veröffentlicht: (2025)
von: Chen, Jingyi, et al.
Veröffentlicht: (2025)
A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems
von: Cuconasu, Florin, et al.
Veröffentlicht: (2024)
von: Cuconasu, Florin, et al.
Veröffentlicht: (2024)
emrQA-msquad: A Medical Dataset Structured with the SQuAD V2.0 Framework, Enriched with emrQA Medical Information
von: Eladio, Jimenez, et al.
Veröffentlicht: (2024)
von: Eladio, Jimenez, et al.
Veröffentlicht: (2024)
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
von: Liu, Naiming, et al.
Veröffentlicht: (2024)
von: Liu, Naiming, et al.
Veröffentlicht: (2024)
Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework
von: Luo, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Luo, Xiaoyu, et al.
Veröffentlicht: (2026)
Evidence Contextualization and Counterfactual Attribution for Conversational QA over Heterogeneous Data with RAG Systems
von: Roy, Rishiraj Saha, et al.
Veröffentlicht: (2024)
von: Roy, Rishiraj Saha, et al.
Veröffentlicht: (2024)
Leveraging Online Data to Enhance Medical Knowledge in a Small Persian Language Model
von: Ghassabi, Mehrdad, et al.
Veröffentlicht: (2025)
von: Ghassabi, Mehrdad, et al.
Veröffentlicht: (2025)
MuCoS: Efficient Drug Target Discovery via Multi Context Aware Sampling in Knowledge Graphs
von: Gul, Haji, et al.
Veröffentlicht: (2025)
von: Gul, Haji, et al.
Veröffentlicht: (2025)
Reverse Image Retrieval Cues Parametric Memory in Multimodal LLMs
von: Xu, Jialiang, et al.
Veröffentlicht: (2024)
von: Xu, Jialiang, et al.
Veröffentlicht: (2024)
DRIV-EX: Counterfactual Explanations for Driving LLMs
von: Cardiel, Amaia, et al.
Veröffentlicht: (2026)
von: Cardiel, Amaia, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scaling Public Health Text Annotation: Zero-Shot Learning vs. Crowdsourcing for Improved Efficiency and Labeling Accuracy
von: Kazari, Kamyar, et al.
Veröffentlicht: (2025) -
Detecting Bias and Enhancing Diagnostic Accuracy in Large Language Models for Healthcare
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2024) -
Can LLMs Explain Themselves Counterfactually?
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2025) -
Aligning (Medical) LLMs for (Counterfactual) Fairness
von: Poulain, Raphael, et al.
Veröffentlicht: (2024) -
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
von: Le, Chenqian, et al.
Veröffentlicht: (2025)