Guardrails for avoiding harmful medical product recommendations and off-label promotion in generative AI models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Lopez-Martinez, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Guardrails for Safe and Secure Healthcare AI
von: Gangavarapu, Ananya
Veröffentlicht: (2024)
von: Gangavarapu, Ananya
Veröffentlicht: (2024)
Building Effective Safety Guardrails in AI Education Tools
von: Clark, Hannah-Beth, et al.
Veröffentlicht: (2025)
von: Clark, Hannah-Beth, et al.
Veröffentlicht: (2025)
A Comparative Evaluation of AI Agent Security Guardrails
von: Li, Qi, et al.
Veröffentlicht: (2026)
von: Li, Qi, et al.
Veröffentlicht: (2026)
First, do no harm: Breaking suicidogenic echo chambers in media recommendation
von: Díaz-Álvarez, Alberto, et al.
Veröffentlicht: (2026)
von: Díaz-Álvarez, Alberto, et al.
Veröffentlicht: (2026)
Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents
von: Kholkar, Gauri, et al.
Veröffentlicht: (2025)
von: Kholkar, Gauri, et al.
Veröffentlicht: (2025)
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
von: She, Yining, et al.
Veröffentlicht: (2025)
von: She, Yining, et al.
Veröffentlicht: (2025)
AI driven health recommender
von: Vignesh, K., et al.
Veröffentlicht: (2024)
von: Vignesh, K., et al.
Veröffentlicht: (2024)
From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
von: Pandya, Ravi, et al.
Veröffentlicht: (2025)
von: Pandya, Ravi, et al.
Veröffentlicht: (2025)
Current state of LLM Risks and AI Guardrails
von: Ayyamperumal, Suriya Ganesh, et al.
Veröffentlicht: (2024)
von: Ayyamperumal, Suriya Ganesh, et al.
Veröffentlicht: (2024)
No Free Lunch with Guardrails
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2025)
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2025)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
von: Jin, Xisen, et al.
Veröffentlicht: (2026)
von: Jin, Xisen, et al.
Veröffentlicht: (2026)
Lattice: Generative Guardrails for Conversational Agents
von: Broadhurst, Emily, et al.
Veröffentlicht: (2026)
von: Broadhurst, Emily, et al.
Veröffentlicht: (2026)
Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers
von: Bertollo, Giacomo, et al.
Veröffentlicht: (2025)
von: Bertollo, Giacomo, et al.
Veröffentlicht: (2025)
Characterizing and modeling harms from interactions with design patterns in AI interfaces
von: Ibrahim, Lujain, et al.
Veröffentlicht: (2024)
von: Ibrahim, Lujain, et al.
Veröffentlicht: (2024)
Challenges in Guardrailing Large Language Models for Science
von: Pantha, Nishan, et al.
Veröffentlicht: (2024)
von: Pantha, Nishan, et al.
Veröffentlicht: (2024)
Provably Secure Agent Guardrail
von: Wu, Benlong, et al.
Veröffentlicht: (2026)
von: Wu, Benlong, et al.
Veröffentlicht: (2026)
Safety Guardrails for LLM-Enabled Robots
von: Ravichandran, Zachary, et al.
Veröffentlicht: (2025)
von: Ravichandran, Zachary, et al.
Veröffentlicht: (2025)
Evaluating adaptive and generative AI-based feedback and recommendations in a knowledge-graph-integrated programming learning system
von: Nongkhai, Lalita Na, et al.
Veröffentlicht: (2026)
von: Nongkhai, Lalita Na, et al.
Veröffentlicht: (2026)
AI Harmonics: a human-centric and harms severity-adaptive AI risk assessment framework
von: Vei, Sofia, et al.
Veröffentlicht: (2025)
von: Vei, Sofia, et al.
Veröffentlicht: (2025)
Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems
von: Lee, Alexander W., et al.
Veröffentlicht: (2025)
von: Lee, Alexander W., et al.
Veröffentlicht: (2025)
Learning Efficient Guardrails for Compliance
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
Building Guardrails for Large Language Models
von: Dong, Yi, et al.
Veröffentlicht: (2024)
von: Dong, Yi, et al.
Veröffentlicht: (2024)
Deep learning with noisy labels in medical prediction problems: a scoping review
von: Wei, Yishu, et al.
Veröffentlicht: (2024)
von: Wei, Yishu, et al.
Veröffentlicht: (2024)
A global log for medical AI
von: Noori, Ayush, et al.
Veröffentlicht: (2025)
von: Noori, Ayush, et al.
Veröffentlicht: (2025)
Behavioral Determinants of Deployed AI Agents in Social Networks: A Multi-Factor Study of Personality, Model, and Guardrail Specification
von: Wilson, Sarah, et al.
Veröffentlicht: (2026)
von: Wilson, Sarah, et al.
Veröffentlicht: (2026)
Generative AI for automatic topic labelling
von: Kozlowski, Diego, et al.
Veröffentlicht: (2024)
von: Kozlowski, Diego, et al.
Veröffentlicht: (2024)
Test-Time Training Undermines Safety Guardrails
von: Antonelli, Simone, et al.
Veröffentlicht: (2026)
von: Antonelli, Simone, et al.
Veröffentlicht: (2026)
A Lightweight Explainable Guardrail for Prompt Safety
von: Islam, Md Asiful, et al.
Veröffentlicht: (2026)
von: Islam, Md Asiful, et al.
Veröffentlicht: (2026)
Generalization in medical AI: a perspective on developing scalable models
von: Zvuloni, Eran, et al.
Veröffentlicht: (2023)
von: Zvuloni, Eran, et al.
Veröffentlicht: (2023)
Climbing the label tree: Hierarchy-preserving contrastive learning for medical imaging
von: Khan, Alif Elham
Veröffentlicht: (2025)
von: Khan, Alif Elham
Veröffentlicht: (2025)
Shift-Up: A Framework for Software Engineering Guardrails in AI-native Software Development -- Initial Findings
von: Lipsanen, Petrus, et al.
Veröffentlicht: (2026)
von: Lipsanen, Petrus, et al.
Veröffentlicht: (2026)
PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
von: Wu, Yaozu, et al.
Veröffentlicht: (2025)
von: Wu, Yaozu, et al.
Veröffentlicht: (2025)
Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences
von: Han, Shanshan, et al.
Veröffentlicht: (2025)
von: Han, Shanshan, et al.
Veröffentlicht: (2025)
AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
von: Luo, Weidi, et al.
Veröffentlicht: (2025)
von: Luo, Weidi, et al.
Veröffentlicht: (2025)
MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support
von: Farinhas, António, et al.
Veröffentlicht: (2026)
von: Farinhas, António, et al.
Veröffentlicht: (2026)
"The Diagram is like Guardrails": Structuring GenAI-assisted Hypotheses Exploration with an Interactive Shared Representation
von: Ding, Zijian, et al.
Veröffentlicht: (2025)
von: Ding, Zijian, et al.
Veröffentlicht: (2025)
Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2026)
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2026)
In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
von: Durner, Nils
Veröffentlicht: (2025)
von: Durner, Nils
Veröffentlicht: (2025)
Trust-Oriented Adaptive Guardrails for Large Language Models
von: Hu, Jinwei, et al.
Veröffentlicht: (2024)
von: Hu, Jinwei, et al.
Veröffentlicht: (2024)
CodeGuard: Improving LLM Guardrails in CS Education
von: Raihan, Nishat, et al.
Veröffentlicht: (2026)
von: Raihan, Nishat, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Enhancing Guardrails for Safe and Secure Healthcare AI
von: Gangavarapu, Ananya
Veröffentlicht: (2024) -
Building Effective Safety Guardrails in AI Education Tools
von: Clark, Hannah-Beth, et al.
Veröffentlicht: (2025) -
A Comparative Evaluation of AI Agent Security Guardrails
von: Li, Qi, et al.
Veröffentlicht: (2026) -
First, do no harm: Breaking suicidogenic echo chambers in media recommendation
von: Díaz-Álvarez, Alberto, et al.
Veröffentlicht: (2026) -
Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents
von: Kholkar, Gauri, et al.
Veröffentlicht: (2025)