Combining Cost-Constrained Runtime Monitors for AI Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Hua, Tim Tian, Baskerville, James, Lemoine, Henri, Hopman, Mia, Bhatt, Aryan, Tracy, Tyler |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Practical Principles for AI Cost and Compute Accounting
by: Casper, Stephen, et al.
Published: (2025)
by: Casper, Stephen, et al.
Published: (2025)
When Should Algorithms Resign? A Proposal for AI Governance
by: Bhatt, Umang, et al.
Published: (2024)
by: Bhatt, Umang, et al.
Published: (2024)
BashArena: A Control Setting for Highly Privileged AI Agents
by: Kaufman, Adam, et al.
Published: (2025)
by: Kaufman, Adam, et al.
Published: (2025)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
by: Dobbe, Roel
Published: (2025)
by: Dobbe, Roel
Published: (2025)
Interoperability in AI Safety Governance: Ethics, Regulations, and Standards
by: Chin, Yik Chan, et al.
Published: (2026)
by: Chin, Yik Chan, et al.
Published: (2026)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
by: Scholefield, Rebecca, et al.
Published: (2025)
by: Scholefield, Rebecca, et al.
Published: (2025)
Safety Cases: A Scalable Approach to Frontier AI Safety
by: Hilton, Benjamin, et al.
Published: (2025)
by: Hilton, Benjamin, et al.
Published: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
by: Clymer, Joshua, et al.
Published: (2024)
by: Clymer, Joshua, et al.
Published: (2024)
Toward an African Agenda for AI Safety
by: Segun, Samuel T., et al.
Published: (2025)
by: Segun, Samuel T., et al.
Published: (2025)
Concrete Problems in AI Safety, Revisited
by: Raji, Inioluwa Deborah, et al.
Published: (2023)
by: Raji, Inioluwa Deborah, et al.
Published: (2023)
The Cost-Benefit of Interdisciplinarity in AI for Mental Health
by: Drakos, Katerina, et al.
Published: (2025)
by: Drakos, Katerina, et al.
Published: (2025)
Runtime Monitoring and Enforcement of Conditional Fairness in Generative AIs
by: Cheng, Chih-Hong, et al.
Published: (2024)
by: Cheng, Chih-Hong, et al.
Published: (2024)
Belief Offloading in Human-AI Interaction
by: Guingrich, Rose E., et al.
Published: (2026)
by: Guingrich, Rose E., et al.
Published: (2026)
Emerging Practices in Frontier AI Safety Frameworks
by: Buhl, Marie Davidsen, et al.
Published: (2025)
by: Buhl, Marie Davidsen, et al.
Published: (2025)
AI Safety: Necessary, but insufficient and possibly problematic
by: P, Deepak
Published: (2024)
by: P, Deepak
Published: (2024)
The Singapore Consensus on Global AI Safety Research Priorities
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
Cryptographic Runtime Governance for Autonomous AI Systems: The Aegis Architecture for Verifiable Policy Enforcement
by: Mazzocchetti, Adam Massimo
Published: (2026)
by: Mazzocchetti, Adam Massimo
Published: (2026)
The Hidden Costs of AI-Mediated Political Outreach: Persuasion and AI Penalties in the US and UK
by: Jungherr, Andreas, et al.
Published: (2026)
by: Jungherr, Andreas, et al.
Published: (2026)
Building Effective Safety Guardrails in AI Education Tools
by: Clark, Hannah-Beth, et al.
Published: (2025)
by: Clark, Hannah-Beth, et al.
Published: (2025)
What Is AI Safety? What Do We Want It to Be?
by: Harding, Jacqueline, et al.
Published: (2025)
by: Harding, Jacqueline, et al.
Published: (2025)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
by: Costa, Mariana Lins
Published: (2026)
by: Costa, Mariana Lins
Published: (2026)
Upstream and Downstream AI Safety: Both on the Same River?
by: McDermid, John, et al.
Published: (2024)
by: McDermid, John, et al.
Published: (2024)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
by: Li, Miles Q., et al.
Published: (2026)
by: Li, Miles Q., et al.
Published: (2026)
Agentic Microphysics: A Manifesto for Generative AI Safety
by: Pierucci, Federico, et al.
Published: (2026)
by: Pierucci, Federico, et al.
Published: (2026)
Probabilistic Analysis of Copyright Disputes and Generative AI Safety
by: Chiba-Okabe, Hiroaki
Published: (2024)
by: Chiba-Okabe, Hiroaki
Published: (2024)
How Hyper-Datafication Impacts the Sustainability Costs in Frontier AI
by: Wilson, Sophia N., et al.
Published: (2026)
by: Wilson, Sophia N., et al.
Published: (2026)
US-China perspectives on extreme AI risks and global governance
by: Wasil, Akash, et al.
Published: (2024)
by: Wasil, Akash, et al.
Published: (2024)
The Global Landscape of Environmental AI Regulation: From the Cost of Reasoning to a Right to Green AI
by: Ebert, Kai, et al.
Published: (2026)
by: Ebert, Kai, et al.
Published: (2026)
Developing Strategies to Increase Capacity in AI Education
by: Cowit, Noah Q., et al.
Published: (2025)
by: Cowit, Noah Q., et al.
Published: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
The Elephant in the Room -- Why AI Safety Demands Diverse Teams
by: Rostcheck, David, et al.
Published: (2024)
by: Rostcheck, David, et al.
Published: (2024)
Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond
by: Han, Shanshan
Published: (2024)
by: Han, Shanshan
Published: (2024)
International Scientific Report on the Safety of Advanced AI (Interim Report)
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
Open Problems in Machine Unlearning for AI Safety
by: Barez, Fazl, et al.
Published: (2025)
by: Barez, Fazl, et al.
Published: (2025)
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
by: Staufer, Leon, et al.
Published: (2026)
by: Staufer, Leon, et al.
Published: (2026)
Know Thyself? On the Incapability and Implications of AI Self-Recognition
by: Bai, Xiaoyan, et al.
Published: (2025)
by: Bai, Xiaoyan, et al.
Published: (2025)
Annotating the Chain-of-Thought: A Behavior-Labeled Dataset for AI Safety
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
Dark Speculation: Combining Qualitative and Quantitative Understanding in Frontier AI Risk Analysis
by: Carpenter, Daniel, et al.
Published: (2025)
by: Carpenter, Daniel, et al.
Published: (2025)
The Hidden AI Race: Tracking Environmental Costs of Innovation
by: Agarwal, Shyam, et al.
Published: (2025)
by: Agarwal, Shyam, et al.
Published: (2025)
Towards interactive evaluations for interaction harms in human-AI systems
by: Ibrahim, Lujain, et al.
Published: (2024)
by: Ibrahim, Lujain, et al.
Published: (2024)
Similar Items
-
Practical Principles for AI Cost and Compute Accounting
by: Casper, Stephen, et al.
Published: (2025) -
When Should Algorithms Resign? A Proposal for AI Governance
by: Bhatt, Umang, et al.
Published: (2024) -
BashArena: A Control Setting for Highly Privileged AI Agents
by: Kaufman, Adam, et al.
Published: (2025) -
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
by: Dobbe, Roel
Published: (2025) -
Interoperability in AI Safety Governance: Ethics, Regulations, and Standards
by: Chin, Yik Chan, et al.
Published: (2026)