Sustaining AI safety: Control-theoretic external impossibility, intrinsic necessity, and structural requirements
Fuente:
arXiv
Guardado en:
| Autor principal: | Mazzu, James M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Supertrust foundational alignment: mutual trust must replace permanent control for safe superintelligence
por: Mazzu, James M.
Publicado: (2024)
por: Mazzu, James M.
Publicado: (2024)
Factorizing formal contexts from closures of necessity operators
por: Aragón, Roberto G., et al.
Publicado: (2026)
por: Aragón, Roberto G., et al.
Publicado: (2026)
Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events
por: Michaelov, James A., et al.
Publicado: (2025)
por: Michaelov, James A., et al.
Publicado: (2025)
Quantifying intrinsic causal contributions via structure preserving interventions
por: Janzing, Dominik, et al.
Publicado: (2020)
por: Janzing, Dominik, et al.
Publicado: (2020)
Evaluating whether AI models would sabotage AI safety research
por: Kirk, Robert, et al.
Publicado: (2026)
por: Kirk, Robert, et al.
Publicado: (2026)
Kallini et al. (2024) do not compare impossible languages with constituency-based ones
por: Hunter, Tim
Publicado: (2024)
por: Hunter, Tim
Publicado: (2024)
Retrying vs Resampling in AI Control
por: Lucassen, James, et al.
Publicado: (2026)
por: Lucassen, James, et al.
Publicado: (2026)
Landscape of AI safety concerns -- A methodology to support safety assurance for AI-based autonomous systems
por: Schnitzer, Ronald, et al.
Publicado: (2024)
por: Schnitzer, Ronald, et al.
Publicado: (2024)
A note on the impossibility of conditional PAC-efficient reasoning in large language models
por: Zeng, Hao
Publicado: (2025)
por: Zeng, Hao
Publicado: (2025)
Comprehensive AI governance requires addressing non-model gains
por: Goemans, Arthur, et al.
Publicado: (2026)
por: Goemans, Arthur, et al.
Publicado: (2026)
Playing games with knowledge: AI-Induced delusions need game theoretic interventions
por: Beaumaster, Will, et al.
Publicado: (2026)
por: Beaumaster, Will, et al.
Publicado: (2026)
Taming the Centaur(s) with LAPITHS: a framework for a theoretically grounded interpretation of AI performances
por: Da Pelo, Matteo, et al.
Publicado: (2026)
por: Da Pelo, Matteo, et al.
Publicado: (2026)
A cross-regional review of AI safety regulations in the commercial aviation
por: Barr, Penny A., et al.
Publicado: (2025)
por: Barr, Penny A., et al.
Publicado: (2025)
Towards evaluations-based safety cases for AI scheming
por: Balesni, Mikita, et al.
Publicado: (2024)
por: Balesni, Mikita, et al.
Publicado: (2024)
Affirmative safety: An approach to risk management for high-risk AI
por: Wasil, Akash R., et al.
Publicado: (2024)
por: Wasil, Akash R., et al.
Publicado: (2024)
Towards provable probabilistic safety for scalable embodied AI systems
por: He, Linxuan, et al.
Publicado: (2025)
por: He, Linxuan, et al.
Publicado: (2025)
A sketch of an AI control safety case
por: Korbak, Tomek, et al.
Publicado: (2025)
por: Korbak, Tomek, et al.
Publicado: (2025)
ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidence
por: Wu, Kevin, et al.
Publicado: (2024)
por: Wu, Kevin, et al.
Publicado: (2024)
The impact of intrinsic rewards on exploration in Reinforcement Learning
por: Kayal, Aya, et al.
Publicado: (2025)
por: Kayal, Aya, et al.
Publicado: (2025)
Sustainable AI Processing at the Edge
por: Ollivier, Sébastien, et al.
Publicado: (2022)
por: Ollivier, Sébastien, et al.
Publicado: (2022)
Super Co-alignment of Human and AI for Sustainable Symbiotic Society
por: Zeng, Yi, et al.
Publicado: (2025)
por: Zeng, Yi, et al.
Publicado: (2025)
"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives
por: Wang, Ding, et al.
Publicado: (2025)
por: Wang, Ding, et al.
Publicado: (2025)
Efficiency Will Not Lead to Sustainable Reasoning AI
por: Wiesner, Philipp, et al.
Publicado: (2025)
por: Wiesner, Philipp, et al.
Publicado: (2025)
On the Sustainability of AI Inferences in the Edge
por: Sobhani, Ghazal, et al.
Publicado: (2025)
por: Sobhani, Ghazal, et al.
Publicado: (2025)
Position: Ensuring mutual privacy is necessary for effective external evaluation of proprietary AI systems
por: Bucknall, Ben, et al.
Publicado: (2025)
por: Bucknall, Ben, et al.
Publicado: (2025)
AI Application in Anti-Money Laundering for Sustainable and Transparent Financial Systems
por: Nie, Chuanhao, et al.
Publicado: (2025)
por: Nie, Chuanhao, et al.
Publicado: (2025)
Quality Assessment of Public Summary of Training Content for GPAI models required by AI Act Article 53(1)(d)
por: Blankvoort, Dick A. H., et al.
Publicado: (2026)
por: Blankvoort, Dick A. H., et al.
Publicado: (2026)
The Environmental Impact of AI Servers and Sustainable Solutions
por: Patel, Aadi, et al.
Publicado: (2025)
por: Patel, Aadi, et al.
Publicado: (2025)
AI Sustainability in Practice Part One: Foundations for Sustainable AI Projects
por: Leslie, David, et al.
Publicado: (2024)
por: Leslie, David, et al.
Publicado: (2024)
AI Sustainability in Practice Part Two: Sustainability Throughout the AI Workflow
por: Leslie, David, et al.
Publicado: (2024)
por: Leslie, David, et al.
Publicado: (2024)
Information-theoretic analysis of world models in optimal reward maximizers
por: Harwood, Alfred, et al.
Publicado: (2026)
por: Harwood, Alfred, et al.
Publicado: (2026)
AI Deception: Risks, Dynamics, and Controls
por: Chen, Boyuan, et al.
Publicado: (2025)
por: Chen, Boyuan, et al.
Publicado: (2025)
An alignment safety case sketch based on debate
por: Buhl, Marie Davidsen, et al.
Publicado: (2025)
por: Buhl, Marie Davidsen, et al.
Publicado: (2025)
The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies
por: Coggins, Sam, et al.
Publicado: (2025)
por: Coggins, Sam, et al.
Publicado: (2025)
Exploring Multimodal Foundation AI and Expert-in-the-Loop for Sustainable Management of Wild Salmon Fisheries in Indigenous Rivers
por: Xu, Chi, et al.
Publicado: (2025)
por: Xu, Chi, et al.
Publicado: (2025)
A Multi-Objective Optimization Approach for Sustainable AI-Driven Entrepreneurship in Resilient Economies
por: ALsobeh, Anas, et al.
Publicado: (2026)
por: ALsobeh, Anas, et al.
Publicado: (2026)
Pareto Optimal Benchmarking of AI Models on ARM Cortex Processors for Sustainable Embedded Systems
por: Jain, Pranay, et al.
Publicado: (2026)
por: Jain, Pranay, et al.
Publicado: (2026)
QAROO: AI-Driven Online Task Offloading for Energy-Efficient and Sustainable MEC Networks
por: Yao, Yongtao, et al.
Publicado: (2026)
por: Yao, Yongtao, et al.
Publicado: (2026)
BashArena: A Control Setting for Highly Privileged AI Agents
por: Kaufman, Adam, et al.
Publicado: (2025)
por: Kaufman, Adam, et al.
Publicado: (2025)
A theoretical guarantee for SyncRank
por: Rao, Yang
Publicado: (2025)
por: Rao, Yang
Publicado: (2025)
Ejemplares similares
-
Supertrust foundational alignment: mutual trust must replace permanent control for safe superintelligence
por: Mazzu, James M.
Publicado: (2024) -
Factorizing formal contexts from closures of necessity operators
por: Aragón, Roberto G., et al.
Publicado: (2026) -
Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events
por: Michaelov, James A., et al.
Publicado: (2025) -
Quantifying intrinsic causal contributions via structure preserving interventions
por: Janzing, Dominik, et al.
Publicado: (2020) -
Evaluating whether AI models would sabotage AI safety research
por: Kirk, Robert, et al.
Publicado: (2026)