Evaluating AI Providers' Frontier Safety Frameworks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Stelling, Lily, Murray, Malcolm, Galizzi, Bruno, Schaffelder, Max, Campos, Siméon, Papadatos, Henry |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Methodology for Quantitative AI Risk Modeling
von: Murray, Malcolm, et al.
Veröffentlicht: (2025)
von: Murray, Malcolm, et al.
Veröffentlicht: (2025)
The Role of Risk Modeling in Advanced AI Risk Management
von: Touzet, Chloé, et al.
Veröffentlicht: (2025)
von: Touzet, Chloé, et al.
Veröffentlicht: (2025)
Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures
von: Stelling, Lily, et al.
Veröffentlicht: (2025)
von: Stelling, Lily, et al.
Veröffentlicht: (2025)
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management
von: Campos, Simeon, et al.
Veröffentlicht: (2025)
von: Campos, Simeon, et al.
Veröffentlicht: (2025)
Lessons from External Review of DeepMind's Scheming Inability Safety Case
von: Barrett, Stephen, et al.
Veröffentlicht: (2026)
von: Barrett, Stephen, et al.
Veröffentlicht: (2026)
Mapping AI Benchmark Data to Quantitative Risk Estimates Through Expert Elicitation
von: Murray, Malcolm, et al.
Veröffentlicht: (2025)
von: Murray, Malcolm, et al.
Veröffentlicht: (2025)
Emerging Practices in Frontier AI Safety Frameworks
von: Buhl, Marie Davidsen, et al.
Veröffentlicht: (2025)
von: Buhl, Marie Davidsen, et al.
Veröffentlicht: (2025)
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
von: Brundage, Miles, et al.
Veröffentlicht: (2026)
von: Brundage, Miles, et al.
Veröffentlicht: (2026)
Toward Quantitative Modeling of Cybersecurity Risks Due to AI Misuse
von: Barrett, Steve, et al.
Veröffentlicht: (2025)
von: Barrett, Steve, et al.
Veröffentlicht: (2025)
The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI Safety
von: Fort, Kristina
Veröffentlicht: (2024)
von: Fort, Kristina
Veröffentlicht: (2024)
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
von: Tong, Haibo, et al.
Veröffentlicht: (2026)
von: Tong, Haibo, et al.
Veröffentlicht: (2026)
Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning
von: Schaffelder, Max, et al.
Veröffentlicht: (2025)
von: Schaffelder, Max, et al.
Veröffentlicht: (2025)
Safety Cases: A Scalable Approach to Frontier AI Safety
von: Hilton, Benjamin, et al.
Veröffentlicht: (2025)
von: Hilton, Benjamin, et al.
Veröffentlicht: (2025)
Enabling Frontier Lab Collaboration to Mitigate AI Safety Risks
von: Felstead, Nicholas
Veröffentlicht: (2025)
von: Felstead, Nicholas
Veröffentlicht: (2025)
Open Problems in Frontier AI Risk Management
von: Ziosi, Marta, et al.
Veröffentlicht: (2026)
von: Ziosi, Marta, et al.
Veröffentlicht: (2026)
Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
Towards Frontier Safety Policies Plus
von: Pistillo, Matteo
Veröffentlicht: (2025)
von: Pistillo, Matteo
Veröffentlicht: (2025)
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
von: Knight, Christina Q., et al.
Veröffentlicht: (2025)
von: Knight, Christina Q., et al.
Veröffentlicht: (2025)
Machine Learning for Public Good: Predicting Urban Crime Patterns to Enhance Community Safety
von: Gupta, Sia, et al.
Veröffentlicht: (2024)
von: Gupta, Sia, et al.
Veröffentlicht: (2024)
Expanding External Access To Frontier AI Models For Dangerous Capability Evaluations
von: Charnock, Jacob, et al.
Veröffentlicht: (2026)
von: Charnock, Jacob, et al.
Veröffentlicht: (2026)
Artificially Fluent: Swahili AI Performance Benchmarks Between English-Trained and Natively-Trained Datasets
von: Jaffer, Sophie, et al.
Veröffentlicht: (2025)
von: Jaffer, Sophie, et al.
Veröffentlicht: (2025)
Vernacularizing Taxonomies of Harm is Essential for Operationalizing Holistic AI Safety
von: Kennedy, Wm. Matthew, et al.
Veröffentlicht: (2024)
von: Kennedy, Wm. Matthew, et al.
Veröffentlicht: (2024)
Evaluating the Goal-Directedness of Large Language Models
von: Everitt, Tom, et al.
Veröffentlicht: (2025)
von: Everitt, Tom, et al.
Veröffentlicht: (2025)
A Grading Rubric for AI Safety Frameworks
von: Alaga, Jide, et al.
Veröffentlicht: (2024)
von: Alaga, Jide, et al.
Veröffentlicht: (2024)
The science and practice of proportionality in AI risk evaluations
von: Mougan, Carlos, et al.
Veröffentlicht: (2026)
von: Mougan, Carlos, et al.
Veröffentlicht: (2026)
Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases
von: Feakins, Shaun, et al.
Veröffentlicht: (2026)
von: Feakins, Shaun, et al.
Veröffentlicht: (2026)
The California Report on Frontier AI Policy
von: Bommasani, Rishi, et al.
Veröffentlicht: (2025)
von: Bommasani, Rishi, et al.
Veröffentlicht: (2025)
AI Safety Evaluations Need To Consider Cascading Effects
von: Neumann, Anna, et al.
Veröffentlicht: (2026)
von: Neumann, Anna, et al.
Veröffentlicht: (2026)
Assessing the Case for Africa-Centric AI Safety Evaluations
von: Ireri, Gathoni, et al.
Veröffentlicht: (2026)
von: Ireri, Gathoni, et al.
Veröffentlicht: (2026)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
von: Ackerman, Gary, et al.
Veröffentlicht: (2025)
von: Ackerman, Gary, et al.
Veröffentlicht: (2025)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
von: Li, Jing-Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2024)
Assurance of Frontier AI Built for National Security
von: Pistillo, Matteo, et al.
Veröffentlicht: (2025)
von: Pistillo, Matteo, et al.
Veröffentlicht: (2025)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)
Institutional AI: A Governance Framework for Distributional AGI Safety
von: Pierucci, Federico, et al.
Veröffentlicht: (2026)
von: Pierucci, Federico, et al.
Veröffentlicht: (2026)
Questionnaire Responses Do not Capture the Safety of AI Agents
von: Hellrigel-Holderbaum, Max, et al.
Veröffentlicht: (2026)
von: Hellrigel-Holderbaum, Max, et al.
Veröffentlicht: (2026)
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
von: Gringras, David, et al.
Veröffentlicht: (2026)
von: Gringras, David, et al.
Veröffentlicht: (2026)
Frontier AI Ethics: Anticipating and Evaluating the Societal Impacts of Language Model Agents
von: Lazar, Seth
Veröffentlicht: (2024)
von: Lazar, Seth
Veröffentlicht: (2024)
Affirmative safety: An approach to risk management for high-risk AI
von: Wasil, Akash R., et al.
Veröffentlicht: (2024)
von: Wasil, Akash R., et al.
Veröffentlicht: (2024)
Frontier AI developers need an internal audit function
von: Schuett, Jonas
Veröffentlicht: (2023)
von: Schuett, Jonas
Veröffentlicht: (2023)
AI Safety Frameworks Should Include Procedures for Model Access Decisions
von: Kembery, Edward, et al.
Veröffentlicht: (2024)
von: Kembery, Edward, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Methodology for Quantitative AI Risk Modeling
von: Murray, Malcolm, et al.
Veröffentlicht: (2025) -
The Role of Risk Modeling in Advanced AI Risk Management
von: Touzet, Chloé, et al.
Veröffentlicht: (2025) -
Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures
von: Stelling, Lily, et al.
Veröffentlicht: (2025) -
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management
von: Campos, Simeon, et al.
Veröffentlicht: (2025) -
Lessons from External Review of DeepMind's Scheming Inability Safety Case
von: Barrett, Stephen, et al.
Veröffentlicht: (2026)