Governable AI: Provable Safety Under Extreme Threat Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Donglin, Liang, Weiyun, Chen, Chunyuan, Xu, Jing, Fu, Yulong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
by: Tong, Haibo, et al.
Published: (2026)
by: Tong, Haibo, et al.
Published: (2026)
Authenticity Debt and the Synthetic Content Threat Landscape: A Layered Framework for Trust, Provenance, and IP Governance in the Generative AI Era
by: Sengupta, Shubhashis, et al.
Published: (2026)
by: Sengupta, Shubhashis, et al.
Published: (2026)
AI-Driven Cyber Threat Intelligence Automation
by: Shah, Shrit, et al.
Published: (2024)
by: Shah, Shrit, et al.
Published: (2024)
AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
by: Lin, Zhiqiang, et al.
Published: (2025)
by: Lin, Zhiqiang, et al.
Published: (2025)
Countering Autonomous Cyber Threats
by: Heckel, Kade M., et al.
Published: (2024)
by: Heckel, Kade M., et al.
Published: (2024)
On Technique Identification and Threat-Actor Attribution using LLMs and Embedding Models
by: Guru, Kyla, et al.
Published: (2025)
by: Guru, Kyla, et al.
Published: (2025)
Future-Back Threat Modeling: A Foresight-Driven Security Framework
by: Van Than, Vu
Published: (2025)
by: Van Than, Vu
Published: (2025)
Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases
by: Feakins, Shaun, et al.
Published: (2026)
by: Feakins, Shaun, et al.
Published: (2026)
Cryptographic Runtime Governance for Autonomous AI Systems: The Aegis Architecture for Verifiable Policy Enforcement
by: Mazzocchetti, Adam Massimo
Published: (2026)
by: Mazzocchetti, Adam Massimo
Published: (2026)
Securing the Future: Proactive Threat Hunting for Sustainable IoT Ecosystems
by: Ghasemshirazi, Saeid, et al.
Published: (2024)
by: Ghasemshirazi, Saeid, et al.
Published: (2024)
Blueprints of Trust: AI System Cards for End to End Transparency and Governance
by: Sidhpurwala, Huzaifa, et al.
Published: (2025)
by: Sidhpurwala, Huzaifa, et al.
Published: (2025)
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
by: Barrett, Anthony M., et al.
Published: (2025)
by: Barrett, Anthony M., et al.
Published: (2025)
AI Agents Under EU Law
by: Nannini, Luca, et al.
Published: (2026)
by: Nannini, Luca, et al.
Published: (2026)
Is Your AI Truly Yours? Leveraging Blockchain for Copyrights, Provenance, and Lineage
by: Wang, Qin, et al.
Published: (2024)
by: Wang, Qin, et al.
Published: (2024)
Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
by: Gressel, Gilad, et al.
Published: (2025)
by: Gressel, Gilad, et al.
Published: (2025)
Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation
by: Heiding, Fred, et al.
Published: (2025)
by: Heiding, Fred, et al.
Published: (2025)
ThreatGPT: An Agentic AI Framework for Enhancing Public Safety through Threat Modeling
by: Zisad, Sharif Noor, et al.
Published: (2025)
by: Zisad, Sharif Noor, et al.
Published: (2025)
Cyber Shadows: Neutralizing Security Threats with AI and Targeted Policy Measures
by: Schmitt, Marc, et al.
Published: (2025)
by: Schmitt, Marc, et al.
Published: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
Phare: A Safety Probe for Large Language Models
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
by: Lin, Justin W., et al.
Published: (2025)
by: Lin, Justin W., et al.
Published: (2025)
RLCP: A Reinforcement Learning-based Copyright Protection Method for Text-to-Image Diffusion Model
by: Shi, Zhuan, et al.
Published: (2024)
by: Shi, Zhuan, et al.
Published: (2024)
The New Frontier of Cybersecurity: Emerging Threats and Innovations
by: Dave, Daksh, et al.
Published: (2023)
by: Dave, Daksh, et al.
Published: (2023)
Private, Verifiable, and Auditable AI Systems
by: South, Tobin
Published: (2025)
by: South, Tobin
Published: (2025)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
AI Propaganda factories with language models
by: Olejnik, Lukasz
Published: (2025)
by: Olejnik, Lukasz
Published: (2025)
Accelerating AI Development with Cyber Arenas
by: Cashman, William, et al.
Published: (2025)
by: Cashman, William, et al.
Published: (2025)
Preserving Decision Sovereignty in Military AI: A Trade-Secret-Safe Architectural Framework for Model Replaceability, Human Authority, and State Control
by: Wei, Peng, et al.
Published: (2026)
by: Wei, Peng, et al.
Published: (2026)
Securing the Future of GenAI: Policy and Technology
by: Christodorescu, Mihai, et al.
Published: (2024)
by: Christodorescu, Mihai, et al.
Published: (2024)
Game Theory Meets LLM and Agentic AI: Reimagining Cybersecurity for the Age of Intelligent Threats
by: Zhu, Quanyan
Published: (2025)
by: Zhu, Quanyan
Published: (2025)
The Pitfalls of "Security by Obscurity" And What They Mean for Transparent AI
by: Hall, Peter, et al.
Published: (2025)
by: Hall, Peter, et al.
Published: (2025)
Black-Box Access is Insufficient for Rigorous AI Audits
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
Coordinated Flaw Disclosure for AI: Beyond Security Vulnerabilities
by: Cattell, Sven, et al.
Published: (2024)
by: Cattell, Sven, et al.
Published: (2024)
Interplay of ISMS and AIMS in context of the EU AI Act
by: Pötsch, Jordan
Published: (2024)
by: Pötsch, Jordan
Published: (2024)
The End Of Universal Lifelong Identifiers: Identity Systems For The AI Era
by: Palakodety, Shriphani
Published: (2025)
by: Palakodety, Shriphani
Published: (2025)
Enabling External Scrutiny of AI Systems with Privacy-Enhancing Technologies
by: Beers, Kendrea, et al.
Published: (2025)
by: Beers, Kendrea, et al.
Published: (2025)
RedTeamLLM: an Agentic AI framework for offensive security
by: Challita, Brian, et al.
Published: (2025)
by: Challita, Brian, et al.
Published: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Naming is framing: How cybersecurity's language problems are repeating in AI governance
by: Potter, Lianne
Published: (2025)
by: Potter, Lianne
Published: (2025)
The Impact of AI on the Cyber Offense-Defense Balance and the Character of Cyber Conflict
by: Lohn, Andrew J.
Published: (2025)
by: Lohn, Andrew J.
Published: (2025)
Similar Items
-
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
by: Tong, Haibo, et al.
Published: (2026) -
Authenticity Debt and the Synthetic Content Threat Landscape: A Layered Framework for Trust, Provenance, and IP Governance in the Generative AI Era
by: Sengupta, Shubhashis, et al.
Published: (2026) -
AI-Driven Cyber Threat Intelligence Automation
by: Shah, Shrit, et al.
Published: (2024) -
AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
by: Lin, Zhiqiang, et al.
Published: (2025) -
Countering Autonomous Cyber Threats
by: Heckel, Kade M., et al.
Published: (2024)