SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866918518960685056 |
|---|---|
| author | Desai, Pratyush Tang, Luoxi Meng, Yuqiao Xi, Zhaohan |
| author_facet | Desai, Pratyush Tang, Luoxi Meng, Yuqiao Xi, Zhaohan |
| contents | Large Language Models (LLMs) are transforming enterprise workflows but introduce security and ethics challenges when employees inadvertently share confidential data or generate policy-violating content. This paper proposes SafeGPT, a two-sided guardrail system preventing sensitive data leakage and unethical outputs. SafeGPT integrates input-side detection/redaction, output-side moderation/reframing, and human-in-the-loop feedback. Experiments demonstrate SafeGPT effectively reduces data leakage risk and biased outputs while maintaining satisfaction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_06366 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use Desai, Pratyush Tang, Luoxi Meng, Yuqiao Xi, Zhaohan Cryptography and Security Artificial Intelligence Large Language Models (LLMs) are transforming enterprise workflows but introduce security and ethics challenges when employees inadvertently share confidential data or generate policy-violating content. This paper proposes SafeGPT, a two-sided guardrail system preventing sensitive data leakage and unethical outputs. SafeGPT integrates input-side detection/redaction, output-side moderation/reframing, and human-in-the-loop feedback. Experiments demonstrate SafeGPT effectively reduces data leakage risk and biased outputs while maintaining satisfaction. |
| title | SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2601.06366 |