SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Desai, Pratyush, Tang, Luoxi, Meng, Yuqiao, Xi, Zhaohan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918518960685056
author Desai, Pratyush
Tang, Luoxi
Meng, Yuqiao
Xi, Zhaohan
author_facet Desai, Pratyush
Tang, Luoxi
Meng, Yuqiao
Xi, Zhaohan
contents Large Language Models (LLMs) are transforming enterprise workflows but introduce security and ethics challenges when employees inadvertently share confidential data or generate policy-violating content. This paper proposes SafeGPT, a two-sided guardrail system preventing sensitive data leakage and unethical outputs. SafeGPT integrates input-side detection/redaction, output-side moderation/reframing, and human-in-the-loop feedback. Experiments demonstrate SafeGPT effectively reduces data leakage risk and biased outputs while maintaining satisfaction.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use
Desai, Pratyush
Tang, Luoxi
Meng, Yuqiao
Xi, Zhaohan
Cryptography and Security
Artificial Intelligence
Large Language Models (LLMs) are transforming enterprise workflows but introduce security and ethics challenges when employees inadvertently share confidential data or generate policy-violating content. This paper proposes SafeGPT, a two-sided guardrail system preventing sensitive data leakage and unethical outputs. SafeGPT integrates input-side detection/redaction, output-side moderation/reframing, and human-in-the-loop feedback. Experiments demonstrate SafeGPT effectively reduces data leakage risk and biased outputs while maintaining satisfaction.
title SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2601.06366