The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making
Fuente:
arXiv
Saved in:
| Main Authors: | Gourabathina, Abinitha, Hao, Yuexing, Gerych, Walter, Ghassemi, Marzyeh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification
by: Hamidieh, Kimia, et al.
Published: (2026)
by: Hamidieh, Kimia, et al.
Published: (2026)
MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations
by: Jin, Qixuan, et al.
Published: (2024)
by: Jin, Qixuan, et al.
Published: (2024)
Robustness Beyond Known Groups with Low-rank Adaptation
by: Gourabathina, Abinitha, et al.
Published: (2026)
by: Gourabathina, Abinitha, et al.
Published: (2026)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
Identifying Implicit Social Biases in Vision-Language Models
by: Hamidieh, Kimia, et al.
Published: (2024)
by: Hamidieh, Kimia, et al.
Published: (2024)
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
by: Gourabathina, Abinitha, et al.
Published: (2026)
by: Gourabathina, Abinitha, et al.
Published: (2026)
Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
by: Javed, Rafiya, et al.
Published: (2025)
by: Javed, Rafiya, et al.
Published: (2025)
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
by: Hao, Yuexing, et al.
Published: (2025)
by: Hao, Yuexing, et al.
Published: (2025)
Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
by: Luo, Jinyuan, et al.
Published: (2025)
by: Luo, Jinyuan, et al.
Published: (2025)
MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
by: Kim, Yubin, et al.
Published: (2024)
by: Kim, Yubin, et al.
Published: (2024)
Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
by: Melo, Rui, et al.
Published: (2025)
by: Melo, Rui, et al.
Published: (2025)
MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations
by: Liu, Genglin, et al.
Published: (2025)
by: Liu, Genglin, et al.
Published: (2025)
An Investigation of Memorization Risk in Healthcare Foundation Models
by: Tonekaboni, Sana, et al.
Published: (2025)
by: Tonekaboni, Sana, et al.
Published: (2025)
The Role of Computing Resources in Publishing Foundation Model Research
by: Hao, Yuexing, et al.
Published: (2025)
by: Hao, Yuexing, et al.
Published: (2025)
FedMedICL: Towards Holistic Evaluation of Distribution Shifts in Federated Medical Imaging
by: Alhamoud, Kumail, et al.
Published: (2024)
by: Alhamoud, Kumail, et al.
Published: (2024)
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
by: Xiao, Yunpeng, et al.
Published: (2025)
by: Xiao, Yunpeng, et al.
Published: (2025)
Rescaling Confidence: What Scale Design Reveals About LLM Metacognition
by: Dai, Yuyang
Published: (2026)
by: Dai, Yuyang
Published: (2026)
Revealing Positive and Negative Role Models to Help People Make Good Decisions
by: Blum, Avrim, et al.
Published: (2026)
by: Blum, Avrim, et al.
Published: (2026)
What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation
by: Gong, Zhuocheng, et al.
Published: (2024)
by: Gong, Zhuocheng, et al.
Published: (2024)
Mining the Mind: What 100M Beliefs Reveal About Frontier LLM Knowledge
by: Ghosh, Shrestha, et al.
Published: (2025)
by: Ghosh, Shrestha, et al.
Published: (2025)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
by: Ye, Chenlu, et al.
Published: (2026)
by: Ye, Chenlu, et al.
Published: (2026)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
by: Xiao, Yuxin, et al.
Published: (2024)
by: Xiao, Yuxin, et al.
Published: (2024)
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
by: Aswal, Darpan, et al.
Published: (2025)
by: Aswal, Darpan, et al.
Published: (2025)
Character-Level Perturbations Disrupt LLM Watermarks
by: Zhang, Zhaoxi, et al.
Published: (2025)
by: Zhang, Zhaoxi, et al.
Published: (2025)
Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses
by: Rupprecht, Jens, et al.
Published: (2025)
by: Rupprecht, Jens, et al.
Published: (2025)
Human Decision-Making with Persuasive and Narrative LLM Explanations
by: Marusich, Laura R., et al.
Published: (2026)
by: Marusich, Laura R., et al.
Published: (2026)
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
by: Wang, Yangyue, et al.
Published: (2026)
by: Wang, Yangyue, et al.
Published: (2026)
PerturbDiff: Functional Diffusion for Single-Cell Perturbation Modeling
by: Yuan, Xinyu, et al.
Published: (2026)
by: Yuan, Xinyu, et al.
Published: (2026)
MedBayes-Lite: Bayesian Uncertainty Quantification for Safe Clinical Decision Support
by: Hossain, Elias, et al.
Published: (2025)
by: Hossain, Elias, et al.
Published: (2025)
Med-CAM: Minimal Evidence for Explaining Medical Decision Making
by: Suhail, Pirzada, et al.
Published: (2026)
by: Suhail, Pirzada, et al.
Published: (2026)
MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning
by: Shen, Yuhao, et al.
Published: (2026)
by: Shen, Yuhao, et al.
Published: (2026)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
by: Puri, Isha, et al.
Published: (2026)
by: Puri, Isha, et al.
Published: (2026)
Position Paper: Post-Solve Robustness in Decision Engines: Feasible Regions and Smoothness Under Perturbations
by: Hu, Yi-Xiang
Published: (2026)
by: Hu, Yi-Xiang
Published: (2026)
ArgMed-Agents: Explainable Clinical Decision Reasoning with LLM Disscusion via Argumentation Schemes
by: Hong, Shengxin, et al.
Published: (2024)
by: Hong, Shengxin, et al.
Published: (2024)
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
by: Chaudhary, Manav, et al.
Published: (2024)
by: Chaudhary, Manav, et al.
Published: (2024)
On The Statistical Representation Properties Of The Perturb-Softmax And The Perturb-Argmax Probability Distributions
by: Indelman, Hedda Cohen, et al.
Published: (2024)
by: Indelman, Hedda Cohen, et al.
Published: (2024)
What Makes a Good Diffusion Planner for Decision Making?
by: Lu, Haofei, et al.
Published: (2025)
by: Lu, Haofei, et al.
Published: (2025)
MedCoAct: Confidence-Aware Multi-Agent Collaboration for Complete Clinical Decision
by: Zheng, Hongjie, et al.
Published: (2025)
by: Zheng, Hongjie, et al.
Published: (2025)
Similar Items
-
Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification
by: Hamidieh, Kimia, et al.
Published: (2026) -
MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations
by: Jin, Qixuan, et al.
Published: (2024) -
Robustness Beyond Known Groups with Low-rank Adaptation
by: Gourabathina, Abinitha, et al.
Published: (2026) -
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
by: Xiao, Yuxin, et al.
Published: (2025) -
Identifying Implicit Social Biases in Vision-Language Models
by: Hamidieh, Kimia, et al.
Published: (2024)