Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
Fuente:
arXiv
Saved in:
| Main Authors: | Mo, Kaijie, Venkatayogi, Siddhartha, Shaib, Chantal, Kouzy, Ramez, Xu, Wei, Wallace, Byron C., Li, Junyi Jessy |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
by: Yun, Hye Sun, et al.
Published: (2026)
by: Yun, Hye Sun, et al.
Published: (2026)
Detection and Measurement of Syntactic Templates in Generated Text
by: Shaib, Chantal, et al.
Published: (2024)
by: Shaib, Chantal, et al.
Published: (2024)
Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
by: Yun, Hye Sun, et al.
Published: (2025)
by: Yun, Hye Sun, et al.
Published: (2025)
QuaLLM-Health: An Adaptation of an LLM-Based Framework for Quantitative Data Extraction from Online Health Discussions
by: Kouzy, Ramez, et al.
Published: (2024)
by: Kouzy, Ramez, et al.
Published: (2024)
Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine
by: Joseph, Sebastian, et al.
Published: (2025)
by: Joseph, Sebastian, et al.
Published: (2025)
Who Taught You That? Tracing Teachers in Model Distillation
by: Wadhwa, Somin, et al.
Published: (2025)
by: Wadhwa, Somin, et al.
Published: (2025)
FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence
by: Joseph, Sebastian Antony, et al.
Published: (2024)
by: Joseph, Sebastian Antony, et al.
Published: (2024)
Measuring AI "Slop" in Text
by: Shaib, Chantal, et al.
Published: (2025)
by: Shaib, Chantal, et al.
Published: (2025)
Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
by: Shaib, Chantal, et al.
Published: (2025)
by: Shaib, Chantal, et al.
Published: (2025)
How Much Annotation is Needed to Compare Summarization Models?
by: Shaib, Chantal, et al.
Published: (2024)
by: Shaib, Chantal, et al.
Published: (2024)
TimeTox: An LLM-Based Pipeline for Automated Extraction of Time Toxicity from Clinical Trial Protocols
by: Vinjamuri, Saketh, et al.
Published: (2026)
by: Vinjamuri, Saketh, et al.
Published: (2026)
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
by: Kambhatla, Gauri, et al.
Published: (2025)
by: Kambhatla, Gauri, et al.
Published: (2025)
Standardizing the Measurement of Text Diversity: A Tool and a Comparative Analysis of Scores
by: Shaib, Chantal, et al.
Published: (2024)
by: Shaib, Chantal, et al.
Published: (2024)
Counterfactual Probing for the Influence of Affect and Specificity on Intergroup Bias
by: Govindarajan, Venkata S, et al.
Published: (2023)
by: Govindarajan, Venkata S, et al.
Published: (2023)
InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification
by: Trienes, Jan, et al.
Published: (2024)
by: Trienes, Jan, et al.
Published: (2024)
Behavioral Analysis of Information Salience in Large Language Models
by: Trienes, Jan, et al.
Published: (2025)
by: Trienes, Jan, et al.
Published: (2025)
Compared to What? Baselines and Metrics for Counterfactual Prompting
by: Yang, Zihao, et al.
Published: (2026)
by: Yang, Zihao, et al.
Published: (2026)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
Counterfactuals As a Means for Evaluating Faithfulness of Attribution Methods in Autoregressive Language Models
by: Kamahi, Sepehr, et al.
Published: (2024)
by: Kamahi, Sepehr, et al.
Published: (2024)
Language Models (Mostly) Do Not Consider Emotion Triggers When Predicting Emotion
by: Singh, Smriti, et al.
Published: (2023)
by: Singh, Smriti, et al.
Published: (2023)
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
by: Hase, Peter, et al.
Published: (2026)
by: Hase, Peter, et al.
Published: (2026)
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
by: Namuduri, Ramya, et al.
Published: (2025)
by: Namuduri, Ramya, et al.
Published: (2025)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
by: Bhan, Milan, et al.
Published: (2025)
by: Bhan, Milan, et al.
Published: (2025)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025)
by: Wadhwa, Manya, et al.
Published: (2025)
Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
by: Ju, Li, et al.
Published: (2026)
by: Ju, Li, et al.
Published: (2026)
WUGNECTIVES: Novel Entity Inferences of Language Models from Discourse Connectives
by: Brubaker, Daniel, et al.
Published: (2025)
by: Brubaker, Daniel, et al.
Published: (2025)
Using Natural Language Explanations to Rescale Human Judgments
by: Wadhwa, Manya, et al.
Published: (2023)
by: Wadhwa, Manya, et al.
Published: (2023)
Learning to Refine with Fine-Grained Natural Language Feedback
by: Wadhwa, Manya, et al.
Published: (2024)
by: Wadhwa, Manya, et al.
Published: (2024)
Strategic Dialogue Assessment: The Crooked Path to Innocence
by: Zheng, Anshun Asher, et al.
Published: (2025)
by: Zheng, Anshun Asher, et al.
Published: (2025)
Diagnosing LLM Reranker Behavior Under Fixed Evidence Pools
by: Arat, Baris, et al.
Published: (2026)
by: Arat, Baris, et al.
Published: (2026)
Wrapper Boxes: Faithful Attribution of Model Predictions to Training Data
by: Su, Yiheng, et al.
Published: (2023)
by: Su, Yiheng, et al.
Published: (2023)
Counterfactual Cultural Cues Reduce Medical QA Accuracy in LLMs: Identifier vs Context Effects
by: Rezaei, Amirhossein Haji Mohammad, et al.
Published: (2026)
by: Rezaei, Amirhossein Haji Mohammad, et al.
Published: (2026)
Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations
by: Hong, Pingjun, et al.
Published: (2026)
by: Hong, Pingjun, et al.
Published: (2026)
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
by: Schimanski, Tobias, et al.
Published: (2024)
by: Schimanski, Tobias, et al.
Published: (2024)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
by: Fayyaz, Mohsen, et al.
Published: (2024)
by: Fayyaz, Mohsen, et al.
Published: (2024)
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
Vector Arithmetic in Concept and Token Subspaces
by: Feucht, Sheridan, et al.
Published: (2025)
by: Feucht, Sheridan, et al.
Published: (2025)
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs
by: Hu, Zichao, et al.
Published: (2024)
by: Hu, Zichao, et al.
Published: (2024)
LLMs Lean on Priors, Not Programming Language Semantics
by: Thimmaiah, Aditya, et al.
Published: (2025)
by: Thimmaiah, Aditya, et al.
Published: (2025)
LinguaSynth: Heterogeneous Linguistic Signals for News Classification
by: Zhang, Duo, et al.
Published: (2025)
by: Zhang, Duo, et al.
Published: (2025)
Similar Items
-
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
by: Yun, Hye Sun, et al.
Published: (2026) -
Detection and Measurement of Syntactic Templates in Generated Text
by: Shaib, Chantal, et al.
Published: (2024) -
Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
by: Yun, Hye Sun, et al.
Published: (2025) -
QuaLLM-Health: An Adaptation of an LLM-Based Framework for Quantitative Data Extraction from Online Health Discussions
by: Kouzy, Ramez, et al.
Published: (2024) -
Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine
by: Joseph, Sebastian, et al.
Published: (2025)