Trust The Typical

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ganguly, Debargha, Sankar, Sreehari, Zhang, Biyao, Singh, Vikash, Gupta, Kanan, Kavuru, Harshini, Luo, Alan, Chen, Weicong, Morningstar, Warren, Machiraju, Raghu, Chaudhary, Vipin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914305572601856
author Ganguly, Debargha
Sankar, Sreehari
Zhang, Biyao
Singh, Vikash
Gupta, Kanan
Kavuru, Harshini
Luo, Alan
Chen, Weicong
Morningstar, Warren
Machiraju, Raghu
Chaudhary, Vipin
author_facet Ganguly, Debargha
Sankar, Sreehari
Zhang, Biyao
Singh, Vikash
Gupta, Kanan
Kavuru, Harshini
Luo, Alan
Chen, Weicong
Morningstar, Warren
Machiraju, Raghu
Chaudhary, Vipin
contents Current approaches to LLM safety fundamentally rely on a brittle cat-and-mouse game of identifying and blocking known threats via guardrails. We argue for a fresh approach: robust safety comes not from enumerating what is harmful, but from deeply understanding what is safe. We introduce Trust The Typical (T3), a framework that operationalizes this principle by treating safety as an out-of-distribution (OOD) detection problem. T3 learns the distribution of acceptable prompts in a semantic space and flags any significant deviation as a potential threat. Unlike prior methods, it requires no training on harmful examples, yet achieves state-of-the-art performance across 18 benchmarks spanning toxicity, hate speech, jailbreaking, multilingual harms, and over-refusal, reducing false positive rates by up to 40x relative to specialized safety models. A single model trained only on safe English text transfers effectively to diverse domains and over 14 languages without retraining. Finally, we demonstrate production readiness by integrating a GPU-optimized version into vLLM, enabling continuous guardrailing during token generation with less than 6% overhead even under dense evaluation intervals on large-scale workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04581
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Trust The Typical
Ganguly, Debargha
Sankar, Sreehari
Zhang, Biyao
Singh, Vikash
Gupta, Kanan
Kavuru, Harshini
Luo, Alan
Chen, Weicong
Morningstar, Warren
Machiraju, Raghu
Chaudhary, Vipin
Computation and Language
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
Current approaches to LLM safety fundamentally rely on a brittle cat-and-mouse game of identifying and blocking known threats via guardrails. We argue for a fresh approach: robust safety comes not from enumerating what is harmful, but from deeply understanding what is safe. We introduce Trust The Typical (T3), a framework that operationalizes this principle by treating safety as an out-of-distribution (OOD) detection problem. T3 learns the distribution of acceptable prompts in a semantic space and flags any significant deviation as a potential threat. Unlike prior methods, it requires no training on harmful examples, yet achieves state-of-the-art performance across 18 benchmarks spanning toxicity, hate speech, jailbreaking, multilingual harms, and over-refusal, reducing false positive rates by up to 40x relative to specialized safety models. A single model trained only on safe English text transfers effectively to diverse domains and over 14 languages without retraining. Finally, we demonstrate production readiness by integrating a GPU-optimized version into vLLM, enabling continuous guardrailing during token generation with less than 6% overhead even under dense evaluation intervals on large-scale workloads.
title Trust The Typical
topic Computation and Language
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2602.04581