Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Han, Shanshan, Avestimehr, Salman, He, Chaoyang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Kick Bad Guys Out! Conditionally Activated Anomaly Detection in Federated Learning with Zero-Knowledge Proof Verification
par: Han, Shanshan, et autres
Publié: (2023)
par: Han, Shanshan, et autres
Publié: (2023)
TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
par: Stripelis, Dimitris, et autres
Publié: (2024)
par: Stripelis, Dimitris, et autres
Publié: (2024)
ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys
par: Niu, Yue, et autres
Publié: (2024)
par: Niu, Yue, et autres
Publié: (2024)
Safety Guardrails for LLM-Enabled Robots
par: Ravichandran, Zachary, et autres
Publié: (2025)
par: Ravichandran, Zachary, et autres
Publié: (2025)
Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond
par: Han, Shanshan
Publié: (2024)
par: Han, Shanshan
Publié: (2024)
Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems
par: Avinash, Karthik, et autres
Publié: (2025)
par: Avinash, Karthik, et autres
Publié: (2025)
Bridging the AI Trustworthiness Gap between Functions and Norms
par: Di Scala, Daan, et autres
Publié: (2025)
par: Di Scala, Daan, et autres
Publié: (2025)
Alopex: A Computational Framework for Enabling On-Device Function Calls with LLMs
par: Ran, Yide, et autres
Publié: (2024)
par: Ran, Yide, et autres
Publié: (2024)
TorchOpera: A Compound AI System for LLM Safety
par: Han, Shanshan, et autres
Publié: (2024)
par: Han, Shanshan, et autres
Publié: (2024)
PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
par: Wu, Yaozu, et autres
Publié: (2025)
par: Wu, Yaozu, et autres
Publié: (2025)
Fox-1: Open Small Language Model for Cloud and Edge
par: Hu, Zijian, et autres
Publié: (2024)
par: Hu, Zijian, et autres
Publié: (2024)
Bridging the Communication Gap: Evaluating AI Labeling Practices for Trustworthy AI Development
par: Fischer, Raphael, et autres
Publié: (2025)
par: Fischer, Raphael, et autres
Publié: (2025)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
par: An, Heajun, et autres
Publié: (2026)
par: An, Heajun, et autres
Publié: (2026)
A Lightweight Explainable Guardrail for Prompt Safety
par: Islam, Md Asiful, et autres
Publié: (2026)
par: Islam, Md Asiful, et autres
Publié: (2026)
Silencing the Guardrails: Inference-Time Jailbreaking via Dynamic Contextual Representation Ablation
par: Xing, Wenpeng, et autres
Publié: (2026)
par: Xing, Wenpeng, et autres
Publié: (2026)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
par: Krishna, Kundan, et autres
Publié: (2025)
par: Krishna, Kundan, et autres
Publié: (2025)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
par: Liu, Zhe, et autres
Publié: (2026)
par: Liu, Zhe, et autres
Publié: (2026)
Test-Time Training Undermines Safety Guardrails
par: Antonelli, Simone, et autres
Publié: (2026)
par: Antonelli, Simone, et autres
Publié: (2026)
AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
par: Luo, Weidi, et autres
Publié: (2025)
par: Luo, Weidi, et autres
Publié: (2025)
ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
par: Feng, Tiantian, et autres
Publié: (2024)
par: Feng, Tiantian, et autres
Publié: (2024)
A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
par: Zhang, Bingjie, et autres
Publié: (2025)
par: Zhang, Bingjie, et autres
Publié: (2025)
Toward Super Agent System with Hybrid AI Routers
par: Yao, Yuhang, et autres
Publié: (2025)
par: Yao, Yuhang, et autres
Publié: (2025)
Building Effective Safety Guardrails in AI Education Tools
par: Clark, Hannah-Beth, et autres
Publié: (2025)
par: Clark, Hannah-Beth, et autres
Publié: (2025)
FedSecurity: Benchmarking Attacks and Defenses in Federated Learning and Federated LLMs
par: Han, Shanshan, et autres
Publié: (2023)
par: Han, Shanshan, et autres
Publié: (2023)
Reconsidering LLM Uncertainty Estimation Methods in the Wild
par: Bakman, Yavuz, et autres
Publié: (2025)
par: Bakman, Yavuz, et autres
Publié: (2025)
Understanding Communication Backends in Cross-Silo Federated Learning
par: Ziashahabi, Amir, et autres
Publié: (2026)
par: Ziashahabi, Amir, et autres
Publié: (2026)
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
par: Huang, Wei-Chieh, et autres
Publié: (2025)
par: Huang, Wei-Chieh, et autres
Publié: (2025)
Clustering and Median Aggregation Improve Differentially Private Inference
par: Amin, Kareem, et autres
Publié: (2025)
par: Amin, Kareem, et autres
Publié: (2025)
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
par: Sreedhar, Makesh Narsimhan, et autres
Publié: (2025)
par: Sreedhar, Makesh Narsimhan, et autres
Publié: (2025)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
par: Ghasemi, Narges, et autres
Publié: (2025)
par: Ghasemi, Narges, et autres
Publié: (2025)
VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration
par: Geng, Jiahui, et autres
Publié: (2025)
par: Geng, Jiahui, et autres
Publié: (2025)
FedGrAINS: Personalized SubGraph Federated Learning with Adaptive Neighbor Sampling
par: Ceyani, Emir, et autres
Publié: (2025)
par: Ceyani, Emir, et autres
Publié: (2025)
Why Do Safety Guardrails Degrade Across Languages?
par: Zhang, Max, et autres
Publié: (2026)
par: Zhang, Max, et autres
Publié: (2026)
Provably Secure Agent Guardrail
par: Wu, Benlong, et autres
Publié: (2026)
par: Wu, Benlong, et autres
Publié: (2026)
Edge Private Graph Neural Networks with Singular Value Perturbation
par: Tang, Tingting, et autres
Publié: (2024)
par: Tang, Tingting, et autres
Publié: (2024)
CodeGuard: Improving LLM Guardrails in CS Education
par: Raihan, Nishat, et autres
Publié: (2026)
par: Raihan, Nishat, et autres
Publié: (2026)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
par: Cho, Dongkyu Derek, et autres
Publié: (2025)
par: Cho, Dongkyu Derek, et autres
Publié: (2025)
Beyond Stars: Bridging the Gap Between Ratings and Review Sentiment with LLM
par: Zuhir, Najla, et autres
Publié: (2025)
par: Zuhir, Najla, et autres
Publié: (2025)
CryptoMamba: Leveraging State Space Models for Accurate Bitcoin Price Prediction
par: Sepehri, Mohammad Shahab, et autres
Publié: (2025)
par: Sepehri, Mohammad Shahab, et autres
Publié: (2025)
OneShield -- the Next Generation of LLM Guardrails
par: DeLuca, Chad, et autres
Publié: (2025)
par: DeLuca, Chad, et autres
Publié: (2025)
Documents similaires
-
Kick Bad Guys Out! Conditionally Activated Anomaly Detection in Federated Learning with Zero-Knowledge Proof Verification
par: Han, Shanshan, et autres
Publié: (2023) -
TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
par: Stripelis, Dimitris, et autres
Publié: (2024) -
ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys
par: Niu, Yue, et autres
Publié: (2024) -
Safety Guardrails for LLM-Enabled Robots
par: Ravichandran, Zachary, et autres
Publié: (2025) -
Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond
par: Han, Shanshan
Publié: (2024)