OneShield -- the Next Generation of LLM Guardrails
Fuente:
arXiv
Salvato in:
| Autori principali: | DeLuca, Chad, Gentile, Anna Lisa, Asthana, Shubhi, Zhang, Bing, Chowdhary, Pawan, Cheng, Kellen, Shbita, Basel, Li, Pengyuan, Ren, Guang-Jie, Gopisetty, Sandeep |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
Backprompting: Leveraging Synthetic Production Data for Health Advice Guardrails
di: Cheng, Kellen Tan, et al.
Pubblicazione: (2025)
di: Cheng, Kellen Tan, et al.
Pubblicazione: (2025)
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
di: Shbita, Basel, et al.
Pubblicazione: (2025)
di: Shbita, Basel, et al.
Pubblicazione: (2025)
WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark from Wikipedia and Wikidata
di: Shbita, Basel, et al.
Pubblicazione: (2026)
di: Shbita, Basel, et al.
Pubblicazione: (2026)
STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
Runtime-Structured Task Decomposition for Agentic Coding Systems
di: Asthana, Shubhi, et al.
Pubblicazione: (2026)
di: Asthana, Shubhi, et al.
Pubblicazione: (2026)
A Systematic Approach for Large Language Models Debugging
di: Shbita, Basel, et al.
Pubblicazione: (2026)
di: Shbita, Basel, et al.
Pubblicazione: (2026)
LLMON: An LLM-native Markup Language to Leverage Structure and Semantics at the LLM Interface
di: Hind, Michael, et al.
Pubblicazione: (2026)
di: Hind, Michael, et al.
Pubblicazione: (2026)
Evaluating Ill-Defined Tasks in Large Language Models
di: Zhou, Yi, et al.
Pubblicazione: (2026)
di: Zhou, Yi, et al.
Pubblicazione: (2026)
LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics
di: Ahmed, Farhan, et al.
Pubblicazione: (2026)
di: Ahmed, Farhan, et al.
Pubblicazione: (2026)
Adaptive PII Mitigation Framework for Large Language Models
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
Adaptive Decoding via Test-Time Policy Learning for Self-Improving Generation
di: Bhardwaj, Asmita, et al.
Pubblicazione: (2026)
di: Bhardwaj, Asmita, et al.
Pubblicazione: (2026)
Enterprise Benchmarks for Large Language Model Evaluation
di: Zhang, Bing, et al.
Pubblicazione: (2024)
di: Zhang, Bing, et al.
Pubblicazione: (2024)
Inked Careers: Tattooing Professional Paths
di: Gabriela DeLuca
Pubblicazione: (2016)
di: Gabriela DeLuca
Pubblicazione: (2016)
Material Selection. B-1 Evaluating and Selecting Learning Materials, Document No. 10d, Revised. Independent Study Training Material for Professional Supervisory Competencies.
di: DeLuca, Joan
Pubblicazione: (1975)
di: DeLuca, Joan
Pubblicazione: (1975)
Letter to the Editor: The Availability of Midwifery Care in Rural United States Communities
di: Myra DeLuca
Pubblicazione: (2025)
di: Myra DeLuca
Pubblicazione: (2025)
Projeto e Metamorfose: Contribuições de Gilberto Velho para os Estudos sobre Carreiras
di: Gabriela DeLuca
Pubblicazione: (2016)
di: Gabriela DeLuca
Pubblicazione: (2016)
Team careers in science: formation, composition and success of persistent collaborations
di: Chowdhary, Sandeep, et al.
Pubblicazione: (2024)
di: Chowdhary, Sandeep, et al.
Pubblicazione: (2024)
On being a forest soil scientist—Reflections at the 14th North American Forest Soils Conference
di: Thomas H. DeLuca
Pubblicazione: (2024)
di: Thomas H. DeLuca
Pubblicazione: (2024)
Aging: An Annotated Guide to Government Publications. The University of Connecticut Library Bibliography Series, Number 3.
di: DeLuca, L., et al.
Pubblicazione: (1975)
di: DeLuca, L., et al.
Pubblicazione: (1975)
STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs
di: An, Sungeun, et al.
Pubblicazione: (2026)
di: An, Sungeun, et al.
Pubblicazione: (2026)
Advancing ARA: the Next-Generation (ARA-Next) DAQ System
di: Giri, Pawan
Pubblicazione: (2025)
di: Giri, Pawan
Pubblicazione: (2025)
Peering Behind the Shield: Guardrail Identification in Large Language Models
di: Yang, Ziqing, et al.
Pubblicazione: (2025)
di: Yang, Ziqing, et al.
Pubblicazione: (2025)
Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers
di: Kezins, Nikita, et al.
Pubblicazione: (2026)
di: Kezins, Nikita, et al.
Pubblicazione: (2026)
PromptShield: Deployable Detection for Prompt Injection Attacks
di: Jacob, Dennis, et al.
Pubblicazione: (2025)
di: Jacob, Dennis, et al.
Pubblicazione: (2025)
The Case for Home-Grown, Sustainable, Next Generation Library Services
di: Haefele, Chad
Pubblicazione: (2011)
di: Haefele, Chad
Pubblicazione: (2011)
Safety Guardrails for LLM-Enabled Robots
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025)
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025)
From challenge to innovation: A grassroots study of teachers’ classroom assessment innovations
di: Christopher DeLuca, et al.
Pubblicazione: (2024)
di: Christopher DeLuca, et al.
Pubblicazione: (2024)
How hermeneutics can guide grading in integrated STEAM education: An evidence‐informed perspective
di: Christopher DeLuca, et al.
Pubblicazione: (2024)
di: Christopher DeLuca, et al.
Pubblicazione: (2024)
Balancing disciplinary and integrated learning: How exemplary STEM teachers negotiate tensions of practice
di: Michelle Dubek, et al.
Pubblicazione: (2024)
di: Michelle Dubek, et al.
Pubblicazione: (2024)
Individual and team performance in cricket
di: Sadekar, Onkar, et al.
Pubblicazione: (2024)
di: Sadekar, Onkar, et al.
Pubblicazione: (2024)
GradShield: Alignment Preserving Finetuning
di: Hu, Zhanhao, et al.
Pubblicazione: (2026)
di: Hu, Zhanhao, et al.
Pubblicazione: (2026)
ShieldNet: Network-Level Guardrails against Emerging Supply-Chain Injections in Agentic Systems
di: Yuan, Zhuowen, et al.
Pubblicazione: (2026)
di: Yuan, Zhuowen, et al.
Pubblicazione: (2026)
Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
di: Liu, Sheng, et al.
Pubblicazione: (2025)
di: Liu, Sheng, et al.
Pubblicazione: (2025)
Current state of LLM Risks and AI Guardrails
di: Ayyamperumal, Suriya Ganesh, et al.
Pubblicazione: (2024)
di: Ayyamperumal, Suriya Ganesh, et al.
Pubblicazione: (2024)
Benchmarking LLM Guardrails in Handling Multilingual Toxicity
di: Yang, Yahan, et al.
Pubblicazione: (2024)
di: Yang, Yahan, et al.
Pubblicazione: (2024)
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
di: Wang, Yan, et al.
Pubblicazione: (2026)
di: Wang, Yan, et al.
Pubblicazione: (2026)
Scientific mobility patterns of Indian researchers: Impact on career growth
di: TM, Siraj, et al.
Pubblicazione: (2025)
di: TM, Siraj, et al.
Pubblicazione: (2025)
Reconsidering the Unheard Screech and The Subaltern in Mahasweta Devi’s Bayen
di: Nusrat Chowdhary
Pubblicazione: (2017)
di: Nusrat Chowdhary
Pubblicazione: (2017)
CodeGuard: Improving LLM Guardrails in CS Education
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
di: Asthana, Shubhi, et al.
Pubblicazione: (2025) -
Backprompting: Leveraging Synthetic Production Data for Health Advice Guardrails
di: Cheng, Kellen Tan, et al.
Pubblicazione: (2025) -
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
di: Shbita, Basel, et al.
Pubblicazione: (2025) -
WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark from Wikipedia and Wikidata
di: Shbita, Basel, et al.
Pubblicazione: (2026) -
STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)