Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Nghiem, Huy, Panda, Swetasudha, Khatwani, Devashish, Nguyen, Huy V., Kenthapadi, Krishnaram, Daumé III, Hal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
by: Nghiem, Huy, et al.
Published: (2026)
by: Nghiem, Huy, et al.
Published: (2026)
'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style
by: Baumler, Connor, et al.
Published: (2026)
by: Baumler, Connor, et al.
Published: (2026)
How May U.S. Courts Scrutinize Their Recidivism Risk Assessment Tools? Contextualizing AI Fairness Criteria on a Judicial Scrutiny-based Framework
by: Nguyen, Tin, et al.
Published: (2025)
by: Nguyen, Tin, et al.
Published: (2025)
When Stereotypes GTG: The Impact of Predictive Text Suggestions on Gender Bias in Human-AI Co-Writing
by: Baumler, Connor, et al.
Published: (2024)
by: Baumler, Connor, et al.
Published: (2024)
From Attribution to Abstention: Training-Free Attention-Based Auditing for Clinical Summarization
by: Yan, Qianqi, et al.
Published: (2026)
by: Yan, Qianqi, et al.
Published: (2026)
Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting Corrections
by: Zhao, Lingjun, et al.
Published: (2024)
by: Zhao, Lingjun, et al.
Published: (2024)
Which Demographic Features Are Relevant for Individual Fairness Evaluation of U.S. Recidivism Risk Assessment Tools?
by: Nguyen, Tin Trung, et al.
Published: (2025)
by: Nguyen, Tin Trung, et al.
Published: (2025)
A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
by: Zhao, Lingjun, et al.
Published: (2025)
by: Zhao, Lingjun, et al.
Published: (2025)
Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
by: Kenthapadi, Krishnaram, et al.
Published: (2024)
by: Kenthapadi, Krishnaram, et al.
Published: (2024)
Steering Safely or Off a Cliff? Rethinking Specificity and Robustness in Inference-Time Interventions
by: Goyal, Navita, et al.
Published: (2026)
by: Goyal, Navita, et al.
Published: (2026)
The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features
by: Goyal, Navita, et al.
Published: (2023)
by: Goyal, Navita, et al.
Published: (2023)
A Framework for LLM-powered Design Assistants
by: Panda, Swaroop
Published: (2025)
by: Panda, Swaroop
Published: (2025)
A Randomized Controlled Trial on Anonymizing Reviewers to Each Other in Peer Review Discussions
by: Rastogi, Charvi, et al.
Published: (2024)
by: Rastogi, Charvi, et al.
Published: (2024)
Effort-aware Fairness: Incorporating a Philosophy-informed, Human-centered Notion of Effort into Algorithmic Fairness Metrics
by: Nguyen, Tin Trung, et al.
Published: (2025)
by: Nguyen, Tin Trung, et al.
Published: (2025)
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
by: Hou, Yu, et al.
Published: (2025)
by: Hou, Yu, et al.
Published: (2025)
AI Safety, Alignment, and Ethics (AI SAE)
by: Waldner, Dylan
Published: (2025)
by: Waldner, Dylan
Published: (2025)
Can Hallucination Correction Improve Video-Language Alignment?
by: Zhao, Lingjun, et al.
Published: (2025)
by: Zhao, Lingjun, et al.
Published: (2025)
Tracer: A Forensic Framework for Detecting Fraudulent Speedruns from Game Replays
by: Yoo, Jaeung Franciskus, et al.
Published: (2025)
by: Yoo, Jaeung Franciskus, et al.
Published: (2025)
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
by: Zhao, Lingjun, et al.
Published: (2026)
by: Zhao, Lingjun, et al.
Published: (2026)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
by: Gor, Maharshi, et al.
Published: (2024)
by: Gor, Maharshi, et al.
Published: (2024)
The Ethics of Advanced AI Assistants
by: Gabriel, Iason, et al.
Published: (2024)
by: Gabriel, Iason, et al.
Published: (2024)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
by: Si, Chenglei, et al.
Published: (2023)
by: Si, Chenglei, et al.
Published: (2023)
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
by: Chehbouni, Khaoula, et al.
Published: (2024)
by: Chehbouni, Khaoula, et al.
Published: (2024)
Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants
by: Reddy, Pavan, et al.
Published: (2025)
by: Reddy, Pavan, et al.
Published: (2025)
"Define Your Terms" : Enhancing Efficient Offensive Speech Classification with Definition
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
Balancing Fairness and Performance in Healthcare AI: A Gradient Reconciliation Approach
by: Wang, Xiaoyang, et al.
Published: (2025)
by: Wang, Xiaoyang, et al.
Published: (2025)
SALAD: Smart AI Language Assistant Daily
by: Nihal, Ragib Amin, et al.
Published: (2024)
by: Nihal, Ragib Amin, et al.
Published: (2024)
Patterns of Student Help-Seeking When Using a Large Language Model-Powered Programming Assistant
by: Sheese, Brad, et al.
Published: (2023)
by: Sheese, Brad, et al.
Published: (2023)
Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
by: Steinacker-Olsztyn, Nicolas, et al.
Published: (2025)
by: Steinacker-Olsztyn, Nicolas, et al.
Published: (2025)
Causal Effect of Group Diversity on Redundancy and Coverage in Peer-Reviewing
by: Goyal, Navita, et al.
Published: (2024)
by: Goyal, Navita, et al.
Published: (2024)
Optimizing Long-Form Clinical Text Generation with Claim-Based Rewards
by: Jhaveri, Samyak, et al.
Published: (2025)
by: Jhaveri, Samyak, et al.
Published: (2025)
Digital Companionship: Overlapping Uses of AI Companions and AI Assistants
by: Manoli, Aikaterina, et al.
Published: (2025)
by: Manoli, Aikaterina, et al.
Published: (2025)
Reinforcing Trustworthiness in Multimodal Emotional Support Systems
by: Le, Huy M., et al.
Published: (2025)
by: Le, Huy M., et al.
Published: (2025)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
by: Zheng, Mingqian, et al.
Published: (2023)
by: Zheng, Mingqian, et al.
Published: (2023)
Enabling the AI Revolution in Healthcare
by: Singh, Mona, et al.
Published: (2025)
by: Singh, Mona, et al.
Published: (2025)
LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion
by: Zhou, Guanghao, et al.
Published: (2026)
by: Zhou, Guanghao, et al.
Published: (2026)
Similar Items
-
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
by: Nghiem, Huy, et al.
Published: (2024) -
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
by: Nghiem, Huy, et al.
Published: (2026) -
'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?
by: Nghiem, Huy, et al.
Published: (2025) -
SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
by: Nghiem, Huy, et al.
Published: (2025) -
"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations
by: Nghiem, Huy, et al.
Published: (2024)