Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Shanshan, Avestimehr, Salman, He, Chaoyang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Kick Bad Guys Out! Conditionally Activated Anomaly Detection in Federated Learning with Zero-Knowledge Proof Verification
di: Han, Shanshan, et al.
Pubblicazione: (2023)
di: Han, Shanshan, et al.
Pubblicazione: (2023)
TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
di: Stripelis, Dimitris, et al.
Pubblicazione: (2024)
di: Stripelis, Dimitris, et al.
Pubblicazione: (2024)
ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys
di: Niu, Yue, et al.
Pubblicazione: (2024)
di: Niu, Yue, et al.
Pubblicazione: (2024)
Safety Guardrails for LLM-Enabled Robots
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025)
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025)
Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond
di: Han, Shanshan
Pubblicazione: (2024)
di: Han, Shanshan
Pubblicazione: (2024)
Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems
di: Avinash, Karthik, et al.
Pubblicazione: (2025)
di: Avinash, Karthik, et al.
Pubblicazione: (2025)
Bridging the AI Trustworthiness Gap between Functions and Norms
di: Di Scala, Daan, et al.
Pubblicazione: (2025)
di: Di Scala, Daan, et al.
Pubblicazione: (2025)
Alopex: A Computational Framework for Enabling On-Device Function Calls with LLMs
di: Ran, Yide, et al.
Pubblicazione: (2024)
di: Ran, Yide, et al.
Pubblicazione: (2024)
TorchOpera: A Compound AI System for LLM Safety
di: Han, Shanshan, et al.
Pubblicazione: (2024)
di: Han, Shanshan, et al.
Pubblicazione: (2024)
PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
di: Wu, Yaozu, et al.
Pubblicazione: (2025)
di: Wu, Yaozu, et al.
Pubblicazione: (2025)
Fox-1: Open Small Language Model for Cloud and Edge
di: Hu, Zijian, et al.
Pubblicazione: (2024)
di: Hu, Zijian, et al.
Pubblicazione: (2024)
Bridging the Communication Gap: Evaluating AI Labeling Practices for Trustworthy AI Development
di: Fischer, Raphael, et al.
Pubblicazione: (2025)
di: Fischer, Raphael, et al.
Pubblicazione: (2025)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
di: An, Heajun, et al.
Pubblicazione: (2026)
di: An, Heajun, et al.
Pubblicazione: (2026)
A Lightweight Explainable Guardrail for Prompt Safety
di: Islam, Md Asiful, et al.
Pubblicazione: (2026)
di: Islam, Md Asiful, et al.
Pubblicazione: (2026)
Silencing the Guardrails: Inference-Time Jailbreaking via Dynamic Contextual Representation Ablation
di: Xing, Wenpeng, et al.
Pubblicazione: (2026)
di: Xing, Wenpeng, et al.
Pubblicazione: (2026)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
di: Krishna, Kundan, et al.
Pubblicazione: (2025)
di: Krishna, Kundan, et al.
Pubblicazione: (2025)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
di: Liu, Zhe, et al.
Pubblicazione: (2026)
di: Liu, Zhe, et al.
Pubblicazione: (2026)
Test-Time Training Undermines Safety Guardrails
di: Antonelli, Simone, et al.
Pubblicazione: (2026)
di: Antonelli, Simone, et al.
Pubblicazione: (2026)
AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
di: Luo, Weidi, et al.
Pubblicazione: (2025)
di: Luo, Weidi, et al.
Pubblicazione: (2025)
ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
di: Zhang, Bingjie, et al.
Pubblicazione: (2025)
di: Zhang, Bingjie, et al.
Pubblicazione: (2025)
Toward Super Agent System with Hybrid AI Routers
di: Yao, Yuhang, et al.
Pubblicazione: (2025)
di: Yao, Yuhang, et al.
Pubblicazione: (2025)
Building Effective Safety Guardrails in AI Education Tools
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025)
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025)
FedSecurity: Benchmarking Attacks and Defenses in Federated Learning and Federated LLMs
di: Han, Shanshan, et al.
Pubblicazione: (2023)
di: Han, Shanshan, et al.
Pubblicazione: (2023)
Reconsidering LLM Uncertainty Estimation Methods in the Wild
di: Bakman, Yavuz, et al.
Pubblicazione: (2025)
di: Bakman, Yavuz, et al.
Pubblicazione: (2025)
Understanding Communication Backends in Cross-Silo Federated Learning
di: Ziashahabi, Amir, et al.
Pubblicazione: (2026)
di: Ziashahabi, Amir, et al.
Pubblicazione: (2026)
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
di: Huang, Wei-Chieh, et al.
Pubblicazione: (2025)
di: Huang, Wei-Chieh, et al.
Pubblicazione: (2025)
Clustering and Median Aggregation Improve Differentially Private Inference
di: Amin, Kareem, et al.
Pubblicazione: (2025)
di: Amin, Kareem, et al.
Pubblicazione: (2025)
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
di: Sreedhar, Makesh Narsimhan, et al.
Pubblicazione: (2025)
di: Sreedhar, Makesh Narsimhan, et al.
Pubblicazione: (2025)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
di: Ghasemi, Narges, et al.
Pubblicazione: (2025)
di: Ghasemi, Narges, et al.
Pubblicazione: (2025)
VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration
di: Geng, Jiahui, et al.
Pubblicazione: (2025)
di: Geng, Jiahui, et al.
Pubblicazione: (2025)
FedGrAINS: Personalized SubGraph Federated Learning with Adaptive Neighbor Sampling
di: Ceyani, Emir, et al.
Pubblicazione: (2025)
di: Ceyani, Emir, et al.
Pubblicazione: (2025)
Why Do Safety Guardrails Degrade Across Languages?
di: Zhang, Max, et al.
Pubblicazione: (2026)
di: Zhang, Max, et al.
Pubblicazione: (2026)
Provably Secure Agent Guardrail
di: Wu, Benlong, et al.
Pubblicazione: (2026)
di: Wu, Benlong, et al.
Pubblicazione: (2026)
Edge Private Graph Neural Networks with Singular Value Perturbation
di: Tang, Tingting, et al.
Pubblicazione: (2024)
di: Tang, Tingting, et al.
Pubblicazione: (2024)
CodeGuard: Improving LLM Guardrails in CS Education
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
di: Cho, Dongkyu Derek, et al.
Pubblicazione: (2025)
di: Cho, Dongkyu Derek, et al.
Pubblicazione: (2025)
Beyond Stars: Bridging the Gap Between Ratings and Review Sentiment with LLM
di: Zuhir, Najla, et al.
Pubblicazione: (2025)
di: Zuhir, Najla, et al.
Pubblicazione: (2025)
CryptoMamba: Leveraging State Space Models for Accurate Bitcoin Price Prediction
di: Sepehri, Mohammad Shahab, et al.
Pubblicazione: (2025)
di: Sepehri, Mohammad Shahab, et al.
Pubblicazione: (2025)
OneShield -- the Next Generation of LLM Guardrails
di: DeLuca, Chad, et al.
Pubblicazione: (2025)
di: DeLuca, Chad, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Kick Bad Guys Out! Conditionally Activated Anomaly Detection in Federated Learning with Zero-Knowledge Proof Verification
di: Han, Shanshan, et al.
Pubblicazione: (2023) -
TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
di: Stripelis, Dimitris, et al.
Pubblicazione: (2024) -
ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys
di: Niu, Yue, et al.
Pubblicazione: (2024) -
Safety Guardrails for LLM-Enabled Robots
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025) -
Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond
di: Han, Shanshan
Pubblicazione: (2024)