Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhu, Wang Bill, Chen, Tianqi, Yu, Xinyan Velocity, Lin, Ching Ying, Law, Jade, Jizzini, Mazen, Nieva, Jorge J., Liu, Ruishan, Jia, Robin |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
par: Sieker, Judith, et autres
Publié: (2025)
par: Sieker, Judith, et autres
Publié: (2025)
Evaluating Large Language Models for Health-related Queries with Presuppositions
par: Kaur, Navreet, et autres
Publié: (2023)
par: Kaur, Navreet, et autres
Publié: (2023)
Evaluating Reasoning Models for Queries with Presuppositions
par: Sathyanathan, Rose, et autres
Publié: (2026)
par: Sathyanathan, Rose, et autres
Publié: (2026)
If We May De-Presuppose: Robustly Verifying Claims through Presupposition-Free Question Decomposition
par: Dipta, Shubhashis Roy, et autres
Publié: (2025)
par: Dipta, Shubhashis Roy, et autres
Publié: (2025)
CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering
par: Li, Yahan, et autres
Publié: (2025)
par: Li, Yahan, et autres
Publié: (2025)
The Presupposition Problem in Representation Genesis
par: Wu, Yiling
Publié: (2026)
par: Wu, Yiling
Publié: (2026)
Beyond Idealized Patients: Evaluating LLMs under Challenging Patient Behaviors in Medical Consultations
par: Li, Yahan, et autres
Publié: (2026)
par: Li, Yahan, et autres
Publié: (2026)
Safer in Translation? Presupposition Robustness in Indic Languages
par: Palnitkar, Aadi, et autres
Publié: (2025)
par: Palnitkar, Aadi, et autres
Publié: (2025)
Generating Complex Code Analyzers from Natural Language Questions
par: Nazari, Amirmohammad, et autres
Publié: (2026)
par: Nazari, Amirmohammad, et autres
Publié: (2026)
PDDL-Mind: Large Language Models are Capable on Belief Reasoning with Reliable State Tracking
par: Zhu, Wang Bill, et autres
Publié: (2026)
par: Zhu, Wang Bill, et autres
Publié: (2026)
EUDAIMONIA: Evaluating Undesirable Dynamics in AI
par: Huang, Jun Rui, et autres
Publié: (2026)
par: Huang, Jun Rui, et autres
Publié: (2026)
PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language Models
par: Zhu, Wang Bill, et autres
Publié: (2025)
par: Zhu, Wang Bill, et autres
Publié: (2025)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
par: Zhu, Yanxu, et autres
Publié: (2024)
par: Zhu, Yanxu, et autres
Publié: (2024)
Let's CONFER: A Dataset for Evaluating Natural Language Inference Models on CONditional InFERence and Presupposition
par: Azin, Tara, et autres
Publié: (2025)
par: Azin, Tara, et autres
Publié: (2025)
Questions, Answers, and Presuppositions
par: Marie Duží
Publié: (2015)
par: Marie Duží
Publié: (2015)
Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs
par: Azin, Tara, et autres
Publié: (2026)
par: Azin, Tara, et autres
Publié: (2026)
What Patients Really Ask: Exploring the Effect of False Assumptions in Patient Information Seeking
par: Xiong, Raymond, et autres
Publié: (2026)
par: Xiong, Raymond, et autres
Publié: (2026)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
par: An, Bang, et autres
Publié: (2024)
par: An, Bang, et autres
Publié: (2024)
Identifying and Answering Questions with False Assumptions: An Interpretable Approach
par: Wang, Zijie, et autres
Publié: (2025)
par: Wang, Zijie, et autres
Publié: (2025)
Syn-QA2: Evaluating False Assumptions in Long-tail Questions with Synthetic QA Datasets
par: Daswani, Ashwin, et autres
Publié: (2024)
par: Daswani, Ashwin, et autres
Publié: (2024)
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
par: Henkel, Owen, et autres
Publié: (2024)
par: Henkel, Owen, et autres
Publié: (2024)
MultiHoax: A Dataset of Multi-hop False-Premise Questions
par: Shafiei, Mohammadamin, et autres
Publié: (2025)
par: Shafiei, Mohammadamin, et autres
Publié: (2025)
Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
par: Wu, Yuanchen, et autres
Publié: (2025)
par: Wu, Yuanchen, et autres
Publié: (2025)
Large Language Models in Fire Engineering: An Examination of Technical Questions Against Domain Knowledge
par: Hostetter, Haley, et autres
Publié: (2024)
par: Hostetter, Haley, et autres
Publié: (2024)
Evaluating Prompting Strategies for Chart Question Answering with Large Language Models
par: Naikar, Ruthuparna, et autres
Publié: (2026)
par: Naikar, Ruthuparna, et autres
Publié: (2026)
Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks
par: Yang, Yuqing, et autres
Publié: (2026)
par: Yang, Yuqing, et autres
Publié: (2026)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
par: Li, Jidong, et autres
Publié: (2025)
par: Li, Jidong, et autres
Publié: (2025)
Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering
par: Jiang, Bowen, et autres
Publié: (2025)
par: Jiang, Bowen, et autres
Publié: (2025)
Predicting Lung Cancer Patient Prognosis with Large Language Models
par: Hu, Danqing, et autres
Publié: (2024)
par: Hu, Danqing, et autres
Publié: (2024)
RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering
par: Qiu, Weikang, et autres
Publié: (2025)
par: Qiu, Weikang, et autres
Publié: (2025)
optipoly: A Python package for boxed-constrained multi-variable polynomial cost functions optimization
par: Alamir, Mazen
Publié: (2024)
par: Alamir, Mazen
Publié: (2024)
MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platform
par: Jung, Hayoung, et autres
Publié: (2025)
par: Jung, Hayoung, et autres
Publié: (2025)
Evaluating the False Trust Engendered by LLM Explanations
par: Palod, Vardhan, et autres
Publié: (2026)
par: Palod, Vardhan, et autres
Publié: (2026)
Art or Artifice? Large Language Models and the False Promise of Creativity
par: Chakrabarty, Tuhin, et autres
Publié: (2023)
par: Chakrabarty, Tuhin, et autres
Publié: (2023)
On the Limitations of Large Language Models (LLMs): False Attribution
par: Adewumi, Tosin, et autres
Publié: (2024)
par: Adewumi, Tosin, et autres
Publié: (2024)
Where do Large Vision-Language Models Look at when Answering Questions?
par: Xing, Xiaoying, et autres
Publié: (2025)
par: Xing, Xiaoying, et autres
Publié: (2025)
LocateBench: Evaluating the Locating Ability of Vision Language Models
par: Chiang, Ting-Rui, et autres
Publié: (2024)
par: Chiang, Ting-Rui, et autres
Publié: (2024)
Type and Complexity Signals in Multilingual Question Representations
par: Kokot, Robin, et autres
Publié: (2025)
par: Kokot, Robin, et autres
Publié: (2025)
Discovering Global False Negatives On the Fly for Self-supervised Contrastive Learning
par: Balmaseda, Vicente, et autres
Publié: (2025)
par: Balmaseda, Vicente, et autres
Publié: (2025)
BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering
par: Chen, Jinghong, et autres
Publié: (2026)
par: Chen, Jinghong, et autres
Publié: (2026)
Documents similaires
-
LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
par: Sieker, Judith, et autres
Publié: (2025) -
Evaluating Large Language Models for Health-related Queries with Presuppositions
par: Kaur, Navreet, et autres
Publié: (2023) -
Evaluating Reasoning Models for Queries with Presuppositions
par: Sathyanathan, Rose, et autres
Publié: (2026) -
If We May De-Presuppose: Robustly Verifying Claims through Presupposition-Free Question Decomposition
par: Dipta, Shubhashis Roy, et autres
Publié: (2025) -
CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering
par: Li, Yahan, et autres
Publié: (2025)