BlueGlass: A Framework for Composite AI Safety
Fuente:
arXiv
Salvato in:
| Autori principali: | Nandigramwar, Harshal, Qutub, Syed, Scholl, Kay-Ulrich |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Global Clipper: Enhancing Safety and Reliability of Transformer-based Object Detection Models
di: Sha, Qutub Syed, et al.
Pubblicazione: (2024)
di: Sha, Qutub Syed, et al.
Pubblicazione: (2024)
Situation Monitor: Diversity-Driven Zero-Shot Out-of-Distribution Detection using Budding Ensemble Architecture for Object Detection
di: Syed, Qutub, et al.
Pubblicazione: (2024)
di: Syed, Qutub, et al.
Pubblicazione: (2024)
Safety is Non-Compositional: A Formal Framework for Capability-Based AI Systems
di: Spera, Cosimo
Pubblicazione: (2026)
di: Spera, Cosimo
Pubblicazione: (2026)
TEAS: Trusted Educational AI Standard: A Framework for Verifiable, Stable, Auditable, and Pedagogically Sound Learning Systems
di: Syed, Abu
Pubblicazione: (2025)
di: Syed, Abu
Pubblicazione: (2025)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
Domain-Agnostic Scalable AI Safety Ensuring Framework
di: Kim, Beomjun, et al.
Pubblicazione: (2025)
di: Kim, Beomjun, et al.
Pubblicazione: (2025)
HySafe-AI: Hybrid Safety Architectural Analysis Framework for AI Systems: A Case Study
di: Pitale, Mandar, et al.
Pubblicazione: (2025)
di: Pitale, Mandar, et al.
Pubblicazione: (2025)
Emerging Practices in Frontier AI Safety Frameworks
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
Leveraging Machine Learning for Early Autism Detection via INDT-ASD Indian Database
di: Shrivastava, Trapti, et al.
Pubblicazione: (2024)
di: Shrivastava, Trapti, et al.
Pubblicazione: (2024)
Cisco Integrated AI Security and Safety Framework Report
di: Chang, Amy, et al.
Pubblicazione: (2025)
di: Chang, Amy, et al.
Pubblicazione: (2025)
DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents
di: Li, Qi, et al.
Pubblicazione: (2025)
di: Li, Qi, et al.
Pubblicazione: (2025)
AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement
di: Zhang, Zhexin, et al.
Pubblicazione: (2025)
di: Zhang, Zhexin, et al.
Pubblicazione: (2025)
AI Safety: A Climb To Armageddon?
di: Cappelen, Herman, et al.
Pubblicazione: (2024)
di: Cappelen, Herman, et al.
Pubblicazione: (2024)
Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks
di: Syed, Toqeer Ali, et al.
Pubblicazione: (2025)
di: Syed, Toqeer Ali, et al.
Pubblicazione: (2025)
A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety
di: François, Camille, et al.
Pubblicazione: (2025)
di: François, Camille, et al.
Pubblicazione: (2025)
Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
di: Grey, Markov, et al.
Pubblicazione: (2025)
di: Grey, Markov, et al.
Pubblicazione: (2025)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
di: Dobbe, Roel
Pubblicazione: (2025)
di: Dobbe, Roel
Pubblicazione: (2025)
Mechanistic Interpretability for AI Safety -- A Review
di: Bereska, Leonard, et al.
Pubblicazione: (2024)
di: Bereska, Leonard, et al.
Pubblicazione: (2024)
A Trilogy of AI Safety Frameworks: Paths from Facts and Knowledge Gaps to Reliable Predictions and New Knowledge
di: Kasif, Simon
Pubblicazione: (2024)
di: Kasif, Simon
Pubblicazione: (2024)
Students' Feedback Requests and Interactions with the SCRIPT Chatbot: Do They Get What They Ask For?
di: Scholl, Andreas, et al.
Pubblicazione: (2025)
di: Scholl, Andreas, et al.
Pubblicazione: (2025)
How Novice Programmers Use and Experience ChatGPT when Solving Programming Exercises in an Introductory Course
di: Scholl, Andreas, et al.
Pubblicazione: (2024)
di: Scholl, Andreas, et al.
Pubblicazione: (2024)
WordAlchemy: A transformer-based Reverse Dictionary
di: Madaswar, Kanhaiya, et al.
Pubblicazione: (2022)
di: Madaswar, Kanhaiya, et al.
Pubblicazione: (2022)
Agentic AI Framework for Smart Inventory Replenishment
di: Syed, Toqeer Ali, et al.
Pubblicazione: (2025)
di: Syed, Toqeer Ali, et al.
Pubblicazione: (2025)
AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries
di: Shu, Wesley, et al.
Pubblicazione: (2026)
di: Shu, Wesley, et al.
Pubblicazione: (2026)
SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification
di: Wang, Hongbo, et al.
Pubblicazione: (2025)
di: Wang, Hongbo, et al.
Pubblicazione: (2025)
NeuroAI for AI Safety
di: Mineault, Patrick, et al.
Pubblicazione: (2024)
di: Mineault, Patrick, et al.
Pubblicazione: (2024)
Epistemic Injustice in Generative AI
di: Kay, Jackie, et al.
Pubblicazione: (2024)
di: Kay, Jackie, et al.
Pubblicazione: (2024)
AI Will Always Love You: Studying Implicit Biases in Romantic AI Companions
di: Grogan, Clare, et al.
Pubblicazione: (2025)
di: Grogan, Clare, et al.
Pubblicazione: (2025)
SafeEmbodAI: a Safety Framework for Mobile Robots in Embodied AI Systems
di: Zhang, Wenxiao, et al.
Pubblicazione: (2024)
di: Zhang, Wenxiao, et al.
Pubblicazione: (2024)
Safety Cases: A Scalable Approach to Frontier AI Safety
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
A Scalable Data-Driven Framework for Systematic Analysis of SEC 10-K Filings Using Large Language Models
di: Daimi, Syed Affan, et al.
Pubblicazione: (2024)
di: Daimi, Syed Affan, et al.
Pubblicazione: (2024)
Agentic AI Framework for Cloudburst Prediction and Coordinated Response
di: Syed, Toqeer Ali, et al.
Pubblicazione: (2025)
di: Syed, Toqeer Ali, et al.
Pubblicazione: (2025)
AI for Service: Proactive Assistance with AI Glasses
di: Wen, Zichen, et al.
Pubblicazione: (2025)
di: Wen, Zichen, et al.
Pubblicazione: (2025)
Governance-Constrained Agentic AI: Blockchain-Enforced Human Oversight for Safety-Critical Wildfire Monitoring
di: Akarma, Ali, et al.
Pubblicazione: (2026)
di: Akarma, Ali, et al.
Pubblicazione: (2026)
Position: AI Safety Requires Effective Controllability
di: Li, Yige, et al.
Pubblicazione: (2026)
di: Li, Yige, et al.
Pubblicazione: (2026)
AI2-Active Safety: AI-enabled Interaction-aware Active Safety Analysis with Vehicle Dynamics
di: Wu, Keshu, et al.
Pubblicazione: (2025)
di: Wu, Keshu, et al.
Pubblicazione: (2025)
Large-Scale Classification of Shortwave Communication Signals with Machine Learning
di: Scholl, Stefan
Pubblicazione: (2025)
di: Scholl, Stefan
Pubblicazione: (2025)
ThreatGPT: An Agentic AI Framework for Enhancing Public Safety through Threat Modeling
di: Zisad, Sharif Noor, et al.
Pubblicazione: (2025)
di: Zisad, Sharif Noor, et al.
Pubblicazione: (2025)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
di: Scholefield, Rebecca, et al.
Pubblicazione: (2025)
di: Scholefield, Rebecca, et al.
Pubblicazione: (2025)
Towards a Comparative Framework for Compositional AI Models
di: Duneau, Tiffany
Pubblicazione: (2025)
di: Duneau, Tiffany
Pubblicazione: (2025)
Documenti analoghi
-
Global Clipper: Enhancing Safety and Reliability of Transformer-based Object Detection Models
di: Sha, Qutub Syed, et al.
Pubblicazione: (2024) -
Situation Monitor: Diversity-Driven Zero-Shot Out-of-Distribution Detection using Budding Ensemble Architecture for Object Detection
di: Syed, Qutub, et al.
Pubblicazione: (2024) -
Safety is Non-Compositional: A Formal Framework for Capability-Based AI Systems
di: Spera, Cosimo
Pubblicazione: (2026) -
TEAS: Trusted Educational AI Standard: A Framework for Verifiable, Stable, Auditable, and Pedagogically Sound Learning Systems
di: Syed, Abu
Pubblicazione: (2025) -
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)