Safety Co-Option and Compromised National Security: The Self-Fulfilling Prophecy of Weakened AI Risk Thresholds
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Khlaaf, Heidy, West, Sarah Myers |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mind the Gap: Foundation Models and the Covert Proliferation of Military Intelligence, Surveillance, and Targeting
par: Khlaaf, Heidy, et autres
Publié: (2024)
par: Khlaaf, Heidy, et autres
Publié: (2024)
LeftoverLocals: Listening to LLM Responses Through Leaked GPU Local Memory
par: Sorensen, Tyler, et autres
Publié: (2024)
par: Sorensen, Tyler, et autres
Publié: (2024)
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
par: Knight, Christina Q., et autres
Publié: (2025)
par: Knight, Christina Q., et autres
Publié: (2025)
Why Agents Compromise Safety Under Pressure
par: Jiang, Hengle, et autres
Publié: (2026)
par: Jiang, Hengle, et autres
Publié: (2026)
Assurance of Frontier AI Built for National Security
par: Pistillo, Matteo, et autres
Publié: (2025)
par: Pistillo, Matteo, et autres
Publié: (2025)
Threshold Crossings as Tail Events for Catastrophic AI Risk
par: Perrier, Elija
Publié: (2025)
par: Perrier, Elija
Publié: (2025)
Astra: AI Safety, Trust, & Risk Assessment
par: Aggarwal, Pranav, et autres
Publié: (2026)
par: Aggarwal, Pranav, et autres
Publié: (2026)
Informing AI Risk Assessment with News Media: Analyzing National and Political Variation in the Coverage of AI Risks
par: Allaham, Mowafak, et autres
Publié: (2025)
par: Allaham, Mowafak, et autres
Publié: (2025)
Enabling Frontier Lab Collaboration to Mitigate AI Safety Risks
par: Felstead, Nicholas
Publié: (2025)
par: Felstead, Nicholas
Publié: (2025)
Building Trust: Foundations of Security, Safety and Transparency in AI
par: Sidhpurwala, Huzaifa, et autres
Publié: (2024)
par: Sidhpurwala, Huzaifa, et autres
Publié: (2024)
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
par: Brundage, Miles, et autres
Publié: (2026)
par: Brundage, Miles, et autres
Publié: (2026)
AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
par: Lin, Zhiqiang, et autres
Publié: (2025)
par: Lin, Zhiqiang, et autres
Publié: (2025)
Patient Safety Risks from AI Scribes: Signals from End-User Feedback
par: Dai, Jessica, et autres
Publié: (2025)
par: Dai, Jessica, et autres
Publié: (2025)
The AI Alignment Paradox
par: West, Robert, et autres
Publié: (2024)
par: West, Robert, et autres
Publié: (2024)
Seeking Human Security Consensus: A Unified Value Scale for Generative AI Value Safety
par: He, Ying, et autres
Publié: (2026)
par: He, Ying, et autres
Publié: (2026)
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
par: Lee, Michael S., et autres
Publié: (2026)
par: Lee, Michael S., et autres
Publié: (2026)
International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications
par: Bengio, Yoshua, et autres
Publié: (2025)
par: Bengio, Yoshua, et autres
Publié: (2025)
Benchmarking and Understanding Safety Risks in AI Character Platforms
par: Wei, Yiluo, et autres
Publié: (2025)
par: Wei, Yiluo, et autres
Publié: (2025)
Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
par: Akiri, Charankumar, et autres
Publié: (2025)
par: Akiri, Charankumar, et autres
Publié: (2025)
Dual-Use AI Face Swap Apps Are Mostly Unsafe: A Systematic Safety Audit
par: Daffalla, Alaa, et autres
Publié: (2026)
par: Daffalla, Alaa, et autres
Publié: (2026)
International AI Safety Report 2025: Second Key Update: Technical Safeguards and Risk Management
par: Bengio, Yoshua, et autres
Publié: (2025)
par: Bengio, Yoshua, et autres
Publié: (2025)
Self-Certification of High-Risk AI Systems: The Example of AI-based Facial Emotion Recognition
par: Autischer, Gregor, et autres
Publié: (2026)
par: Autischer, Gregor, et autres
Publié: (2026)
Universal Safety Controllers with Learned Prophecies
par: Finkbeiner, Bernd, et autres
Publié: (2025)
par: Finkbeiner, Bernd, et autres
Publié: (2025)
AI Safety for Everyone
par: Gyevnar, Balint, et autres
Publié: (2025)
par: Gyevnar, Balint, et autres
Publié: (2025)
The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI Safety
par: Fort, Kristina
Publié: (2024)
par: Fort, Kristina
Publié: (2024)
Safety First: Psychological Safety as the Key to AI Transformation
par: Reich, Aaron, et autres
Publié: (2026)
par: Reich, Aaron, et autres
Publié: (2026)
How Should AI Safety Benchmarks Benchmark Safety?
par: Yu, Cheng, et autres
Publié: (2026)
par: Yu, Cheng, et autres
Publié: (2026)
An Empirical Analysis on the Use and Reporting of National Security Letters
par: Bellon, Alex, et autres
Publié: (2024)
par: Bellon, Alex, et autres
Publié: (2024)
AI Safety, Alignment, and Ethics (AI SAE)
par: Waldner, Dylan
Publié: (2025)
par: Waldner, Dylan
Publié: (2025)
Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures
par: Stelling, Lily, et autres
Publié: (2025)
par: Stelling, Lily, et autres
Publié: (2025)
Safety Features for a Centralised AGI Project
par: Hastings-Woodhouse, Sarah
Publié: (2025)
par: Hastings-Woodhouse, Sarah
Publié: (2025)
Safety cases for frontier AI
par: Buhl, Marie Davidsen, et autres
Publié: (2024)
par: Buhl, Marie Davidsen, et autres
Publié: (2024)
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
par: Vaugrante, Laurène, et autres
Publié: (2025)
par: Vaugrante, Laurène, et autres
Publié: (2025)
Training Compute Thresholds: Features and Functions in AI Regulation
par: Heim, Lennart, et autres
Publié: (2024)
par: Heim, Lennart, et autres
Publié: (2024)
Latent Profiles of AI Risk Perception and Their Differential Association with Community Driving Safety Concerns: A Person-Centered Analysis
par: Rafe, Amir, et autres
Publié: (2026)
par: Rafe, Amir, et autres
Publié: (2026)
In Quest of an Extensible Multi-Level Harm Taxonomy for Adversarial AI: Heart of Security, Ethical Risk Scoring and Resilience Analytics
par: Khan, Javed I., et autres
Publié: (2026)
par: Khan, Javed I., et autres
Publié: (2026)
Generative AI in Saudi Arabia: A National Survey of Adoption, Risks, and Public Perceptions
par: AlDakheel, Abdulaziz, et autres
Publié: (2026)
par: AlDakheel, Abdulaziz, et autres
Publié: (2026)
The BIG Argument for AI Safety Cases
par: Habli, Ibrahim, et autres
Publié: (2025)
par: Habli, Ibrahim, et autres
Publié: (2025)
Persuasion and Safety in the Era of Generative AI
par: Kong, Haein
Publié: (2025)
par: Kong, Haein
Publié: (2025)
International AI Safety Report 2026
par: Bengio, Yoshua, et autres
Publié: (2026)
par: Bengio, Yoshua, et autres
Publié: (2026)
Documents similaires
-
Mind the Gap: Foundation Models and the Covert Proliferation of Military Intelligence, Surveillance, and Targeting
par: Khlaaf, Heidy, et autres
Publié: (2024) -
LeftoverLocals: Listening to LLM Responses Through Leaked GPU Local Memory
par: Sorensen, Tyler, et autres
Publié: (2024) -
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
par: Knight, Christina Q., et autres
Publié: (2025) -
Why Agents Compromise Safety Under Pressure
par: Jiang, Hengle, et autres
Publié: (2026) -
Assurance of Frontier AI Built for National Security
par: Pistillo, Matteo, et autres
Publié: (2025)