Saved in:
| Main Authors: | Yun, Hye Sun, Zhang, Karen Y. C., Kouzy, Ramez, Marshall, Iain J., Li, Junyi Jessy, Wallace, Byron C. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.07963 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
by: Yun, Hye Sun, et al.
Published: (2026)
by: Yun, Hye Sun, et al.
Published: (2026)
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
by: Mo, Kaijie, et al.
Published: (2026)
by: Mo, Kaijie, et al.
Published: (2026)
Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine
by: Joseph, Sebastian, et al.
Published: (2025)
by: Joseph, Sebastian, et al.
Published: (2025)
Automatically Extracting Numerical Results from Randomized Controlled Trials with Large Language Models
by: Yun, Hye Sun, et al.
Published: (2024)
by: Yun, Hye Sun, et al.
Published: (2024)
QuaLLM-Health: An Adaptation of an LLM-Based Framework for Quantitative Data Extraction from Online Health Discussions
by: Kouzy, Ramez, et al.
Published: (2024)
by: Kouzy, Ramez, et al.
Published: (2024)
Do Multi-Document Summarization Models Synthesize?
by: DeYoung, Jay, et al.
Published: (2023)
by: DeYoung, Jay, et al.
Published: (2023)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
Detection and Measurement of Syntactic Templates in Generated Text
by: Shaib, Chantal, et al.
Published: (2024)
by: Shaib, Chantal, et al.
Published: (2024)
TimeTox: An LLM-Based Pipeline for Automated Extraction of Time Toxicity from Clinical Trial Protocols
by: Vinjamuri, Saketh, et al.
Published: (2026)
by: Vinjamuri, Saketh, et al.
Published: (2026)
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs
by: Hu, Zichao, et al.
Published: (2024)
by: Hu, Zichao, et al.
Published: (2024)
LLMs Lean on Priors, Not Programming Language Semantics
by: Thimmaiah, Aditya, et al.
Published: (2025)
by: Thimmaiah, Aditya, et al.
Published: (2025)
Do LLMs Benefit From Their Own Words?
by: Huang, Jenny Y., et al.
Published: (2026)
by: Huang, Jenny Y., et al.
Published: (2026)
FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence
by: Joseph, Sebastian Antony, et al.
Published: (2024)
by: Joseph, Sebastian Antony, et al.
Published: (2024)
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs
by: Sheffield, William, et al.
Published: (2025)
by: Sheffield, William, et al.
Published: (2025)
Question answering systems for health professionals at the point of care -- a systematic review
by: Kell, Gregory, et al.
Published: (2024)
by: Kell, Gregory, et al.
Published: (2024)
Large Language Models Produce Responses Perceived to be Empathic
by: Lee, Yoon Kyung, et al.
Published: (2024)
by: Lee, Yoon Kyung, et al.
Published: (2024)
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification
by: Trienes, Jan, et al.
Published: (2024)
by: Trienes, Jan, et al.
Published: (2024)
Do Large Language Models Get Caught in Hofstadter-Mobius Loops?
by: Hryszko, Jaroslaw
Published: (2026)
by: Hryszko, Jaroslaw
Published: (2026)
Discourse Diversity in Multi-Turn Empathic Dialogue
by: Zhan, Hongli, et al.
Published: (2026)
by: Zhan, Hongli, et al.
Published: (2026)
Language Models (Mostly) Do Not Consider Emotion Triggers When Predicting Emotion
by: Singh, Smriti, et al.
Published: (2023)
by: Singh, Smriti, et al.
Published: (2023)
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
SPRI: Aligning Large Language Models with Context-Situated Principles
by: Zhan, Hongli, et al.
Published: (2025)
by: Zhan, Hongli, et al.
Published: (2025)
Open (Clinical) LLMs are Sensitive to Instruction Phrasings
by: Arroyo, Alberto Mario Ceballos, et al.
Published: (2024)
by: Arroyo, Alberto Mario Ceballos, et al.
Published: (2024)
Do they mean 'us'? Interpreting Referring Expressions in Intergroup Bias
by: Govindarajan, Venkata S, et al.
Published: (2024)
by: Govindarajan, Venkata S, et al.
Published: (2024)
The Dual-Route Model of Induction
by: Feucht, Sheridan, et al.
Published: (2025)
by: Feucht, Sheridan, et al.
Published: (2025)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025)
by: Wadhwa, Manya, et al.
Published: (2025)
WebWalker: Benchmarking LLMs in Web Traversal
by: Wu, Jialong, et al.
Published: (2025)
by: Wu, Jialong, et al.
Published: (2025)
CREATE: Testing LLMs for Associative Creativity
by: Wadhwa, Manya, et al.
Published: (2026)
by: Wadhwa, Manya, et al.
Published: (2026)
Do Large Language Models Understand Word Senses?
by: Meconi, Domenico, et al.
Published: (2025)
by: Meconi, Domenico, et al.
Published: (2025)
MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
From Tokens to Words: On the Inner Lexicon of LLMs
by: Kaplan, Guy, et al.
Published: (2024)
by: Kaplan, Guy, et al.
Published: (2024)
Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
by: Li, Junyi Jessy, et al.
Published: (2026)
by: Li, Junyi Jessy, et al.
Published: (2026)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
by: Dai, Xunlian, et al.
Published: (2025)
by: Dai, Xunlian, et al.
Published: (2025)
Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
Leveraging ChatGPT in Pharmacovigilance Event Extraction: An Empirical Study
by: Sun, Zhaoyue, et al.
Published: (2024)
by: Sun, Zhaoyue, et al.
Published: (2024)
Do Activation Verbalization Methods Convey Privileged Information?
by: Li, Millicent, et al.
Published: (2025)
by: Li, Millicent, et al.
Published: (2025)
On-the-fly Definition Augmentation of LLMs for Biomedical NER
by: Munnangi, Monica, et al.
Published: (2024)
by: Munnangi, Monica, et al.
Published: (2024)
Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
Circuit Distillation
by: Wadhwa, Somin, et al.
Published: (2025)
by: Wadhwa, Somin, et al.
Published: (2025)
Similar Items
-
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
by: Yun, Hye Sun, et al.
Published: (2026) -
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
by: Mo, Kaijie, et al.
Published: (2026) -
Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine
by: Joseph, Sebastian, et al.
Published: (2025) -
Automatically Extracting Numerical Results from Randomized Controlled Trials with Large Language Models
by: Yun, Hye Sun, et al.
Published: (2024) -
QuaLLM-Health: An Adaptation of an LLM-Based Framework for Quantitative Data Extraction from Online Health Discussions
by: Kouzy, Ramez, et al.
Published: (2024)