Reducing Large Language Model Safety Risks in Women's Health using Semantic Entropy
Fuente:
arXiv
Saved in:
| Main Authors: | Penny-Dimri, Jahan C., Bachmann, Magdalena, Cooke, William R., Mathewlynn, Sam, Dockree, Samuel, Tolladay, John, Kossen, Jannik, Li, Lin, Gal, Yarin, Jones, Gabriel Davis |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
by: Nikitin, Alexander, et al.
Published: (2024)
by: Nikitin, Alexander, et al.
Published: (2024)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
by: Kossen, Jannik, et al.
Published: (2023)
by: Kossen, Jannik, et al.
Published: (2023)
Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
by: Tjandra, Benedict Aaron, et al.
Published: (2024)
by: Tjandra, Benedict Aaron, et al.
Published: (2024)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
by: Kossen, Jannik, et al.
Published: (2024)
by: Kossen, Jannik, et al.
Published: (2024)
The Benefits and Risks of Transductive Approaches for AI Fairness
by: Razzak, Muhammed, et al.
Published: (2024)
by: Razzak, Muhammed, et al.
Published: (2024)
Gaussian See, Gaussian Do: Semantic 3D Motion Transfer from Multiview Video
by: Bekor, Yarin, et al.
Published: (2025)
by: Bekor, Yarin, et al.
Published: (2025)
Scaling Up Active Testing to Large Language Models
by: Berrada, Gabrielle, et al.
Published: (2025)
by: Berrada, Gabrielle, et al.
Published: (2025)
Language Models Change Facts Based on the Way You Talk
by: Kearney, Matthew, et al.
Published: (2025)
by: Kearney, Matthew, et al.
Published: (2025)
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025)
by: Schut, Lisa, et al.
Published: (2025)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
Uncertainty-Aware Step-wise Verification with Generative Reward Models
by: Ye, Zihuiwen, et al.
Published: (2025)
by: Ye, Zihuiwen, et al.
Published: (2025)
The Latency Wall: Benchmarking Off-the-Shelf Emotion Recognition for Real-Time Virtual Avatars
by: Benyamin, Yarin
Published: (2026)
by: Benyamin, Yarin
Published: (2026)
Deep Bayesian Active Learning for Preference Modeling in Large Language Models
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
Explaining Explainability: Recommendations for Effective Use of Concept Activation Vectors
by: Nicolson, Angus, et al.
Published: (2024)
by: Nicolson, Angus, et al.
Published: (2024)
Maximizing Phylogenetic Diversity under Time Pressure: Planning with Extinctions Ahead
by: Jones, Mark, et al.
Published: (2024)
by: Jones, Mark, et al.
Published: (2024)
Estimating the Hallucination Rate of Generative AI
by: Jesson, Andrew, et al.
Published: (2024)
by: Jesson, Andrew, et al.
Published: (2024)
Predicting Fetal Outcomes from Cardiotocography Signals Using a Supervised Variational Autoencoder
by: Tolladay, John, et al.
Published: (2025)
by: Tolladay, John, et al.
Published: (2025)
TextCAVs: Debugging vision models using text
by: Nicolson, Angus, et al.
Published: (2024)
by: Nicolson, Angus, et al.
Published: (2024)
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
by: Yadav, Karmesh, et al.
Published: (2025)
by: Yadav, Karmesh, et al.
Published: (2025)
Fine-tuning can cripple your foundation model; preserving features may be the solution
by: Mukhoti, Jishnu, et al.
Published: (2023)
by: Mukhoti, Jishnu, et al.
Published: (2023)
Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare
by: Shankar, Ravi, et al.
Published: (2025)
by: Shankar, Ravi, et al.
Published: (2025)
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
by: Gupta, Gunshi, et al.
Published: (2025)
by: Gupta, Gunshi, et al.
Published: (2025)
The Curse of Recursion: Training on Generated Data Makes Models Forget
by: Shumailov, Ilia, et al.
Published: (2023)
by: Shumailov, Ilia, et al.
Published: (2023)
First‐trimester biomarkers of gestational diabetes mellitus: A scoping review
by: May Swinburne, et al.
Published: (2025)
by: May Swinburne, et al.
Published: (2025)
Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability
by: Lia, Nusrat Jahan, et al.
Published: (2026)
by: Lia, Nusrat Jahan, et al.
Published: (2026)
Improving Engagement and Efficacy of mHealth Micro-Interventions for Stress Coping: an In-The-Wild Study
by: Yehuda, Chaya Ben, et al.
Published: (2024)
by: Yehuda, Chaya Ben, et al.
Published: (2024)
Bayesian Preference Elicitation with Language Models
by: Handa, Kunal, et al.
Published: (2024)
by: Handa, Kunal, et al.
Published: (2024)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026)
by: Ye, Zihuiwen, et al.
Published: (2026)
Towards a Neural Debugger for Python
by: Beck, Maximilian, et al.
Published: (2026)
by: Beck, Maximilian, et al.
Published: (2026)
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
by: Gupta, Gunshi, et al.
Published: (2024)
by: Gupta, Gunshi, et al.
Published: (2024)
Candidate Monotonicity and Proportionality for Lotteries and Non-Resolute Rules
by: Peters, Jannik
Published: (2024)
by: Peters, Jannik
Published: (2024)
Large Language Models and Video Games: A Preliminary Scoping Review
by: Sweetser, Penny
Published: (2024)
by: Sweetser, Penny
Published: (2024)
A Custom-Built Ambient Scribe Reduces Cognitive Load and Documentation Burden for Telehealth Clinicians
by: Morse, Justin, et al.
Published: (2025)
by: Morse, Justin, et al.
Published: (2025)
Hybrid Physics-Machine Learning Models for Quantitative Electron Diffraction Refinements
by: Malik, Shreshth A., et al.
Published: (2025)
by: Malik, Shreshth A., et al.
Published: (2025)
Proportional Fairness in Clustering: A Social Choice Perspective
by: Kellerhals, Leon, et al.
Published: (2023)
by: Kellerhals, Leon, et al.
Published: (2023)
Iterative Deployment Improves Planning Skills in LLMs
by: Corrêa, Augusto B., et al.
Published: (2025)
by: Corrêa, Augusto B., et al.
Published: (2025)
Generating Robot Constitutions & Benchmarks for Semantic Safety
by: Sermanet, Pierre, et al.
Published: (2025)
by: Sermanet, Pierre, et al.
Published: (2025)
Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content
by: Aperstein, Yehudit, et al.
Published: (2025)
by: Aperstein, Yehudit, et al.
Published: (2025)
Evaluating the Clinical Safety of LLMs in Response to High-Risk Mental Health Disclosures
by: Shah, Siddharth, et al.
Published: (2025)
by: Shah, Siddharth, et al.
Published: (2025)
Similar Items
-
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
by: Nikitin, Alexander, et al.
Published: (2024) -
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
by: Kossen, Jannik, et al.
Published: (2023) -
Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
by: Tjandra, Benedict Aaron, et al.
Published: (2024) -
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
by: Kossen, Jannik, et al.
Published: (2024) -
The Benefits and Risks of Transductive Approaches for AI Fairness
by: Razzak, Muhammed, et al.
Published: (2024)